Skip to content

Publish evidence-led convergence practice and retrieval comparisons - #2

Merged
seathatflowsinourveins merged 2 commits into
mainfrom
codex/evidence-convergence-practice
Sep 20, 2026
Merged

seathatflowsinourveins merged 2 commits into
mainfrom
codex/evidence-convergence-practice

Conversation

@seathatflowsinourveins

@seathatflowsinourveins seathatflowsinourveins commented Sep 20, 2026 •

Copy link
Copy Markdown
Owner

This adds a reusable research-to-acceptance practice and two executed retrieval comparisons to the native stack. Public repository, star and curated-list discovery now connects to pinned source review, frozen experiments, observed results and explicit adoption gates across six capability areas.

The source wave refreshes one public owned repository and all 342 public stars, inspects selected links in three pinned awesome lists, and reviews seven candidates. Registered decisions extend the typed union to 504 repository identities. Primary source and license hashes are retained separately from execution evidence.

The exact-release Agent Retrieval Bench replay covers all 101 trace2code cases with zero skips. Its lexical baseline has higher mean Recall@20 than BM25 on that subset; a separate frozen 24-question public-document fixture favors BM25 at Recall@1 on its 20 answerable cases. The four no-answer cases retain rankings without inventing an abstention score. Both packets include exact pins, hashes, per-case results, replay instructions and limits; they establish retrieval behavior rather than agent repair, model savings or host-wide adoption.

Validation:

  • Full repository suite: 273 passed, 38 optional SDK skips, zero failures/errors.
  • All publication/catalog/index validators and diff checks pass; 434 files are integrity-listed.
  • Fifteen upstream corpus/baseline tests pass at the exact ARB release commit.
  • Independent reviews verified source/input/artifact hashes and recomputed external metrics. An independent execution of the documented local recipe reproduced all 48 rankings and metrics.
  • Redacted secret scan passes. Published inputs are public metadata and public-source-derived fixtures.

The change targets current main, preserving its Mac worker fix and newer security-identity research. All original convergence source and result artifacts remain unchanged.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-20T02:28:26.623149Z d8bd981 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@seathatflowsinourveins
seathatflowsinourveins changed the base branch from codex/mac-worker-timeout-current to main September 20, 2026 02:25
@seathatflowsinourveins
seathatflowsinourveins merged commit 88fe380 into main Sep 20, 2026
4 checks passed
seathatflowsinourveins added a commit that referenced this pull request Sep 23, 2026
…ng for bootstrap

Codex's round-3e review of c18aea3 found two new Highs and four Mediums --
each earlier round's hand-written rollback had closed one edge case and
opened another (an unowned rm -rf, a partial hard-copy backup, a reload
skipped at a step boundary, cleanup left armed after success, a signal
racing a foreground mv, an unguarded rm aborting recovery under set -e).
Per the coordinator's explicit direction, this round changes the approach
instead of patching further: simplify toward idempotent convergence, the
way `brew services` works, with step tables and signal-injection tests
proving every row.

Both scripts now explicitly trap INT, TERM and HUP (`exit 130/143/129`
only; recovery logic stays in the EXIT handler alone) -- bash only
reliably defers a caught signal to a completed step when the signal is
explicitly trapped, not left at its default disposition (confirmed
directly: an untrapped SIGINT sent to a bash script with only an EXIT trap
registered does not interrupt it at all in this environment, while an
untrapped SIGTERM does -- the asymmetry Codex's finding 5 was pointing at).

bootstrap-macos.sh (High #1, unowned rm -rf): new prune_old_version()
deletes a superseded versioned directory only when its canonical form is a
direct child of the canonical tools/ directory with a name matching
<id>-<version>-<stamp>; an external target, a relative "../" target, a
non-matching name, or a symlink loop (canonical_path's python3
os.path.realpath handles cycles without hanging) is logged and left in
place, never deleted. Pruning is now genuinely best-effort: a rejected or
failed deletion warns and never fails the install. Four new tests, one per
adversarial target shape, confirmed to fail against c18aea3 (the function
does not exist there).

launchd-agents.sh -- full redesign, net roughly flat in size (+22 lines)
despite ~390 changed, since one pure function replaces four tracking
markers (pending_backup_plist/_dest/_reload_needed, pending_dest_tmp) and
their two independent, historically drifting recovery implementations:
- backup_plist is a hard link (`ln`, High #2), not `cp`: it either exists
  completely or not at all, so a partial copy can never overwrite the
  intact original. Same directory as dest_plist (round 3d), never this
  script's own state_dir.
- Staging (cp to dest_plist.new, then mv over dest_plist) is unchanged in
  shape but the name lost its PID suffix, matching the backup's own
  naming, so a stale artifact from a crashed run is visible to and
  reconciled by the very next run, not just the same process.
- reconcile_install replaces every marker: it reads dest_plist's identity
  against backup_plist (-ef, since the backup is a hard link) and
  launchd's own current loaded state, and converges to whichever ONE
  action that state calls for (drop the backup, reload, restore-and-
  reload, or -- if nothing converges -- keep the backup and name the exact
  re-run command). It is the ONLY recovery logic: called explicitly after
  a normal attempt, and identically as the EXIT trap for anything that
  interrupts one. `set +e` inside it, explicit checks throughout, per the
  coordinator's instruction that the handler must never itself become
  fatal under the script's own `set -e`.
- record_enabled_label now runs as soon as install actually commits to a
  label (not only on success), so a label that ends up needing attention
  is never re-refused as unowned on retry; a was_already_enabled snapshot,
  taken before that point, keeps the L2 bootout gate correctly checking
  ownership as it stood BEFORE this run, not after.

New tests: a shared _signal_shim helper plus a parameterized step x signal
matrix (ln and mv, TERM and INT) proving retry-converges for launchd; a
dedicated SIGTERM/SIGINT-parameterized migration test and three more for
bootstrap-macos.sh's prune_old_version. Every new/changed test confirmed
against c18aea3 by stashing the fix; several rewritten round-3b/c/d tests
(now asserting the new design's own honest behavior -- e.g. a fresh
install that never loads is left on disk with a manual-retry message,
never silently deleted) checked to still exercise the coordinator's
original findings under the new mechanism.

Full acceptance re-run on this commit: full suite (2821 tests) passes
except the one evidence-registration lag scripts/validate.py itself always
flags immediately after an edit (resolved by the registration commit that
follows); named macOS/launchd modules (126 tests) and the bash-3.2/
launchd-gated subset (56 tests) under a real compiled GNU bash 3.2.0 all
pass; --plan clean under both bash versions; shellcheck and zizmor clean;
git diff --check clean; guarded gitleaks clean. origin/main had not moved
since it was already merged into c18aea3's own parent, so no new merge was
needed this round. No macOS host ran any of this.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 23, 2026
…hd agents (brew-services semantics), embedding acceptance, recording-tooling CI (#94)

* Pin socraticode and platform-dependency integrity, add launchd agents for the macOS adoption profile

adoption/bootstrap-macos.sh: replace the jq-only Homebrew block with a
bash-3.2-safe loop over the six floating formulae (jq, python@3.13, ripgrep,
coreutils, restic, shellcheck), installing only what `brew list --versions`
reports missing, gated on --skip-system-packages/--plan, and echoed in
--plan. socraticode now has a reviewed npm pin (--ignore-scripts, matching
the Linux recipe) instead of being a documented, unpinned skip, so
documented_unpinned_ids is now empty (guarded for bash 3.2). Add
verify_platform_dependency: a fail-closed check, run right after `npm
install`, that the darwin-arm64 platform_dependency npm resolves for codex
and claude-code (name/version/integrity) matches the host's own
.package-lock.json or package.json _integrity.

adoption/pins-macos-arm64.json: add the socraticode pin (sha256 verified by
downloading the tarball and cross-checking against `npm view
socraticode@1.14.0 dist.integrity`) and a platform_dependency entry for each
of @openai/codex and @anthropic-ai/claude-code, with npm dist.integrity read
2026-09-23. @openai/codex-darwin-arm64 is not itself a published package
name; `npm view @openai/codex@0.155.1 optionalDependencies` shows it as an
npm: alias to @openai/codex@0.155.1-darwin-arm64, recorded as such.

New adoption/launchd/: three plist templates (qdrant, ai-memory, the
llama.cpp Metal embedding server on port 8232) using the same ${NAME}
placeholder convention as adoption/templates/, plus launchd-agents.sh
(render/lint/install/status/remove; plutil -lint with a plistlib fallback;
launchctl bootstrap/print/bootout; remove only boots out a label this
script's own state file recorded as enabled, never deleting data). New
tools/adoption/render_launchd.py reuses render_config.py's render_one/
load_host_values/parse_set_values. New adoption/hosts/macos-example.json
supplies the render_config placeholder schema with /Users/example paths and
cites hardware-profiles.json's macos-arm64-48gb-projected id; EMBED_URL is
the bare host:port form (not the http://-prefixed form first drafted) since
project.codex.config.template.toml already prepends http:// for it.

.github/workflows/adoption-bootstrap.yml: add tools/adoption/** to the path
filters; add a launchd render/lint/bootstrap-qdrant/health-check/bootout
step to the existing bootstrap-macos job; add bootstrap-macos-brew (runs the
script without --skip-system-packages, asserting python3.13/rg/restic/
shellcheck/gsha256sum) and validate-macos (the repository's own validators,
the adoption test modules as a gate, and one full python3 -m unittest
recorded as a non-gating artifact), both on macos-15.

adoption/platforms/macos-arm64.md: prerequisites, exit-code table,
darwin-arm64 pin table and launchd section updated to match; the
"unpinned"/documented-skip wording for socraticode is removed since it is
now pinned.

tests/test_adoption_bootstrap_macos.py: update DOCUMENTED_SKIPS, the brew
block assertions, and add PlatformDependencyVerificationTests (offline,
fixture pins + fake .package-lock.json/npm shim; fixture integrity values
are assembled from concatenated parts, not written as single literals).
New tests/test_adoption_launchd.py covers template rendering/schema,
render_launchd.py, launchd-agents.sh structure and an offline
render/lint/install/status/remove cycle against a fake launchctl.

Evidence class: no macOS host ran anything here. Everything is either local
integration (Linux-side unit tests, shellcheck, plistlib parsing) or a
hosted macos-15 runner the coordinator dispatches later.
adoption/platforms/macos-arm64.md stays drafted_not_accepted, and
adoption/manifest.json is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Re-hash the evidence registry and rebuild the ecosystem explorer

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix round: real npm/lockfile evidence, bash 3.2 consumption bug, launchd ownership, XML escaping, harden-runner

Addresses the Opus evidence review (fail), the Codex cross-family review
(fail), and hosted macos-15 CI run 35812470345 on PR #94, which reproduced
three of these findings directly and one more (git_revision) on its own.

1. npm platform-dependency verification never could have passed: measured
   directly against npm 11.19.0 on this project's own host (existing
   ecosystem prefixes, e.g. mcporter-0.13.13, and a fresh `npm install
   --global --prefix <dir> is-odd`), a --global --prefix install writes no
   lockfile anywhere under the prefix, no package.json _integrity, and
   places a platform-specific optional dependency NESTED under the parent
   (lib/node_modules/<parent>/node_modules/<platform-dep>), not at its own
   top-level name -- reproduced here with two real npm-packed local
   packages (tests.RealNpmLockfileEvidenceTests), not asserted from memory.
   verify_platform_dependency (post-hoc lockfile/package.json read) is
   replaced by install_platform_dependency: fetch the platform tarball
   itself (dist.tarball, added to each platform_dependency pin alongside a
   freshly re-hashed sha256), verify it with the same fetch() every other
   pin uses, then `npm install --global --prefix <dir> <alias>@file:<path>`
   the already-verified local archive -- confirmed with real npm
   (PlatformDependencyInstallTests) that this places the verified bytes at
   the alias name regardless of the tarball's own internal package.json
   name, which is what Node's require() actually resolves by.

2. bootstrap-macos.sh:192's `for allowed_id in "${allowed_unpinned_ids[@]}"`
   was unguarded: the array's assignment was bash-3.2-safe but this later
   consumption of it was not, and it aborts with "allowed_unpinned_ids[@]:
   unbound variable" whenever it is genuinely empty -- exactly the hosted
   CI failure. Fixed with the same ${arr[@]+"${arr[@]}"} guard used
   elsewhere. This fix round compiled a real bash 3.2.0 from source
   (GNU bash source, ftp.gnu.org) to reproduce the abort and confirm the
   fix dynamically -- not just structurally, which is all bash 5's relaxed
   nounset handling of empty arrays can ever show. ScriptBehaviorUnderRealBash32Tests
   runs the exact --plan and exit-3 paths under that real 3.2 binary when
   BASH32_BINARY is set, or on a real Mac (/bin/bash genuinely is 3.2 there).

2b. tests/test_adoption_launchd.py's end-to-end cycle test asserted the
   plistlib fallback text unconditionally, which fails wherever plutil is
   present (a real Mac; not this dev host). Relaxed to a generic success
   check, with two new dedicated, PATH-controlled tests exercising each
   path deliberately (test_lint_falls_back_to_plistlib_when_plutil_is_absent,
   test_lint_prefers_plutil_when_present), so the behavior is proven either
   way rather than incidentally on whichever host happens to run it.

2c. scripts/adoption_status.py's git_revision(root) compared git's own
   resolved --show-toplevel path against an unresolved caller root; macOS
   routes its default tempdir through /var -> /private/var (reproduced here
   with a plain symlink on Linux), so a direct caller with an unresolved
   root -- exactly what the existing unit test does -- got None instead of
   the real SHA on a real Mac. inspect_adoption already resolves root
   first, so this was latent; git_revision now resolves its own root
   parameter, closing the gap for every caller.

3. launchd-agents.sh install now refuses to overwrite an existing
   ~/Library/LaunchAgents/<label>.plist that this script did not itself
   record as enabled, and creates state/logs and state/qdrant (qdrant's new
   WorkingDirectory) before installing -- launchd does not create missing
   parent directories for StandardOutPath/StandardErrorPath. remove keeps
   ownership (does not forget the label) when launchctl bootout fails, and
   on success now deletes the copied plist under ~/Library/LaunchAgents
   (not just booting it out) so RunAtLoad cannot reload it at the next
   login; it still never touches the component's own data or logs.
   llama-embed is left out of the default install/status/remove set until a
   model argument exists (render/lint still cover it, and an explicit
   --label still installs it).

4. tools/adoption/render_launchd.py rendered plists via a raw
   string.Template substitution on template text, which is not
   XML-escaping: a substituted value containing "&" (a real path such as
   /Volumes/R&D/eco) would corrupt the document. Rendering now parses the
   template as a plist first (plistlib tolerates the still-literal ${NAME}
   placeholders as ordinary string content), substitutes inside every
   string leaf of the resulting structure, and re-serializes with
   plistlib.dumps, so plistlib's own escaping covers every value; proven
   with a real "&"-containing host value round-tripped through valid XML.

5. The bootstrap-macos-brew and validate-macos jobs, and the existing
   bootstrap-macos job (which had none), now start with harden-runner in
   audit mode: the pinned version's own README (read at its exact pinned
   commit) states macOS runners get full audit-mode support, which is the
   only mode this project ever uses. tests/test_workflow_hardening.py's
   blanket "macOS is exempt" rule is narrowed to a named,
   ownership-scoped exemption for hardware-profile-smoke.yml only (owned by
   a different, concurrent task and off limits to edit here), with a new
   regression test enumerating this branch's three macOS jobs explicitly.

6. bootstrap-macos-brew no longer runs actions/setup-python (which could
   shadow python@3.13 with a GitHub-provided interpreter on PATH, making
   the check pass without ever exercising Homebrew's own install); it now
   checks a short, explicit list of brew --prefix candidates for
   python3.13 and fails with a clear message (not a false pass) if none
   resolve.

7. Path filters gained tests/test_adoption_*.py and the five scripts
   validate-macos runs.

8. The launchd smoke step's "qdrant binary absent" branch exited 0; qdrant
   is a pinned, always-selected component (not itself in required_commands,
   which is owned by adoption/manifest.json and not edited here), so this
   now exits 1 -- its absence after a successful bootstrap is a real
   failure to surface.

9. adoption/platforms/macos-arm64.md's two previously-cited green hosted
   runs are now marked as predating this fix round, naming the actual
   failing run (35812470345) and what it caught, so a reader does not treat
   stale hosted evidence as covering the current script.

10. brew_formulae's install order is now jq alone (needed by the
   --profile/pins validation that follows), then that validation, then the
   remaining five formulae -- previously all six installed before any
   validation ran.

Evidence class: still no macOS host ran anything here. The npm/lockfile
finding and the bash 3.2 abort are now independently reproduced (real npm,
a real compiled bash 3.2.0, and a real /var-style symlink) rather than
argued from memory; the launchd ownership/XML-escaping/harden-runner fixes
remain offline local_integration and structural evidence pending the next
hosted macos-15 run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Rehash the evidence registry via host_receipts.register_file and rebuild the explorer

Used scripts/host_receipts.py's register_file (an absolute Path.cwd() root,
per the coordinator's fix-round instructions) for every file this fix round
changed, instead of hand-patching manifests/evidence.json: the six already
registered in the previous round's rehash (adoption-bootstrap.yml,
bootstrap-macos.sh, pins-macos-arm64.json, macos-arm64.md,
test_adoption_bootstrap_macos.py) plus scripts/adoption_status.py, and this
round's four new files that were not previously registered
(com.native-stack.qdrant.plist.template, launchd-agents.sh,
test_adoption_launchd.py, test_workflow_hardening.py, render_launchd.py).
register_file rewrites the whole manifest via json.dumps(indent=2), so this
diff also reformats entries this round did not otherwise touch; the content
change is exactly the updated/added hash and byte-count records.

docs/ecosystem/index.html: regenerated with `scripts/build_ecosystem.py
--write` from the manifests above and registered the same way; --check
passes.

python3 scripts/validate.py, scripts/validate_catalogs.py,
scripts/validate_foundation.py --root . --json, scripts/landscape.py --root
., scripts/build_ecosystem.py --check and
tools/sota-convergence/build_verdicts.py --check all pass after this commit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Round 2: defeat the npm shadow-copy, launchd orphan/unloaded/directory fixes, remaining lows

Addresses the Opus round-2 re-review (fail): one High, one Medium, several
Lows. Full evidence for each below.

**High: npm's own unverified fetch shadowed the verified platform tarball.**
install_npm installed the wrapper without deferring its own optional-dependency
resolution, so npm auto-fetched the darwin-arm64 platform package itself,
unverified, and nested it under the wrapper's own node_modules -- and Node's
require.resolve checks that nested copy before a top-level sibling, so the
previous round's `<alias>@file:<path>` top-level install was never what
codex's bin/codex.js or claude-code's install.cjs actually resolved.
Measured directly (npm 11.19.0) that neither --omit=optional, --no-optional
nor NPM_CONFIG_OMIT=optional suppress the physical on-disk fetch for this
install shape (a single global tarball install, no lockfile), and that
pre-placing verified content at the nested path before installing the
wrapper does not survive it either (npm overwrote it, "changed N packages").
install_platform_dependency now: installs the wrapper with --ignore-scripts
first (deferring its lifecycle scripts); asks Node itself, via
require.resolve scoped to the wrapper's own directory, exactly where it
would resolve the dependency from; deletes that path and extracts the
independently SHA-256-verified tarball there (falling back to a top-level
alias only if Node found nothing to shadow); re-verifies resolution and the
resolved version against the pin, fail closed on either mismatch; then
install_npm runs `npm rebuild --global` for the wrapper, npm's own
documented way to run the lifecycle scripts an --ignore-scripts install
deferred, now against the verified copy. Proven end to end with real npm
against two fixtures mimicking the two real shapes: a codex-like wrapper
(no lifecycle scripts, resolves at every invocation) and a claude-code-like
one (a postinstall that copies from wherever require.resolve finds the
platform package into its own bin/ file, exactly like install.cjs) --
tests.PlatformDependencyInstallTests.test_codex_like_wrapper_resolves_the_verified_dependency_not_the_shadowed_one
and .test_claude_code_like_wrapper_postinstall_copies_the_verified_dependency,
each asserting the FINAL resolved/copied content is the verified fixture
payload, never the differently-content "unverified" one an attacker-controlled
registry fetch stands in for.

**Medium: launchd-agents.sh install left an orphaned, unowned plist on a
failed bootstrap.** cp ran before launchctl bootstrap; a failed bootstrap
left the copied plist under ~/Library/LaunchAgents where RunAtLoad would
reload it at the next login, and neither a retry (the ownership check) nor
remove (never touches an unrecorded label) could ever reach it again. Now
rolls the copy back (rm -f the destination) and exits 1 on a failed
bootstrap, never recording it as enabled.

**Lows:**
- remove on an owned-but-unloaded label always failed calling bootout on
  something not loaded; now checks `launchctl print` first and, if not
  loaded, cleans up (deletes the plist, forgets the label) directly instead.
- install's directory creation was hardcoded to this shell's own
  ECO_INSTALL_ROOT (state/logs, state/qdrant), which breaks a plist rendered
  with --host against a different host's ECO_ROOT value. Now reads each
  rendered plist's own StandardOutPath/StandardErrorPath/WorkingDirectory
  (via a new plist_value helper, plutil-or-plistlib-fallback like cmd_lint)
  and creates directories from those declared paths instead.
- tests/test_workflow_hardening.py's MACOS_JOBS_OWNED_ELSEWHERE exempted the
  whole hardware-profile-smoke.yml file; keyed as
  "hardware-profile-smoke.yml:macos-profile" now, so that file's own ubuntu
  linux-profile job stays checked (new regression test confirms it does).
- Path filters gained scripts/adoption_status.py in both push/pull_request
  lists.
- render_launchd.py's render_plist now also catches ValueError from
  string.Template.substitute (a malformed placeholder, e.g. a bare trailing
  "$") and raises it as the same RenderError every other failure uses,
  instead of an uncaught traceback.
- scripts/adoption_status.py's git_revision now also catches RuntimeError,
  which Path.resolve() raises on Python before 3.13 for a symlink loop
  (3.13+ instead returns the unresolved remainder); reproduced with a real
  two-node symlink loop on this host's Python 3.12.

Also fixed in passing: tests/test_adoption_bootstrap.py's
test_workflow_installs_the_manifest_supported_python_before_status did a
whole-file text.index("scripts/adoption_status.py") search, which this
round's new path-filter entry (a bare, unquoted-command list item earlier in
the file than any job's real invocation) made match before the intended
"actions/setup-python@" occurrence; narrowed the search to the literal
"python3 scripts/adoption_status.py" invocation.

Evidence class: the npm shadow-copy fix and the symlink-loop fix are real
npm/Python behavior, measured and reproduced directly on this host (no
macOS host ran any of this); the launchd fixes remain offline
local_integration and structural evidence pending the next hosted macos-15
run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Round 3: reinstall-while-loaded, arbitrary-path deletion, npmrc/postinstall gaps, atomic staging

Addresses the Opus pass-with-findings review of round 2 (1c01f4f): the High
(npm shadow-copy) is confirmed closed by that review (real tarballs
downloaded, all four hashes confirmed, tarball contents and npm rebuild
scope checked, prepare confirmed not to run). This commit fixes the
remaining Medium/Low/Nit findings.

1. Medium, launchd-agents.sh install: re-running install on a label that is
   owned AND already loaded (e.g. after a re-render) had cp overwrite the
   live plist in place, then launchctl bootstrap fail (already bootstrapped),
   then the rollback delete the just-overwritten file out from under the
   still-running service. install now checks launchctl print first and
   boots the label out before reinstalling if it is loaded (refusing and
   leaving it running if that bootout itself fails); it only ever deletes
   the destination on a failed bootstrap when it did not already exist,
   otherwise it backs it up first and restores (and best-effort reloads) it.
   New shim test with a stateful launchctl (bootstrap sets a per-label
   loaded marker, print reads it, bootout clears it) drives install into
   exactly this already-loaded-reinstall-then-bootstrap-fails path and
   checks the original plist survives byte for byte.

2. Medium, bootstrap-macos.sh install_platform_dependency: `rm -rf` ran on
   whatever path require.resolve returned. Measured directly: Node's
   require.resolve, even given an explicit `paths` array, still searches its
   GLOBAL_FOLDERS fallback (NODE_PATH entries, $HOME/.node_modules, etc. --
   Node's own module docs); setting NODE_PATH to a decoy directory made it
   resolve OUTSIDE the prefix entirely, and the previous code would have
   deleted that decoy -- a real, exploitable arbitrary-path deletion, not a
   theoretical one. The resolved path is now accepted only when it equals
   the exact expected nested path
   ($prefix/lib/node_modules/$package/node_modules/$dep_name/package.json);
   anything else falls back to the existing top-level alias instead of
   touching whatever require.resolve actually named. After installing the
   verified copy, the script now also asserts the resolved path equals
   $target_dir/package.json and that the installed package's own `name`
   field equals the pin's `resolved_package` (pinned since round 1, unread
   by any check until now), fail closed on either mismatch. New tests: a
   real decoy directory via NODE_PATH that survives completely untouched
   while the verified content still lands via the top-level fallback, and a
   resolved_package mismatch that fails closed even though the sha256 and
   version both verify.

3. Low: npm rebuild now passes --ignore-scripts=false explicitly, since a
   user-level .npmrc with ignore-scripts=true would otherwise make it return
   early without running anything. Added `postinstall_binary_check` (an
   optional pin field: platform_file / wrapper_file) and a cmp -s comparison
   after npm rebuild for claude-code specifically (its install.cjs copies
   package/claude, confirmed by downloading and `tar -tzf`ing the real
   2.1.278 darwin-arm64 tarball, into bin/claude.exe): the check fails
   closed if what actually landed in bin/ does not byte-for-byte match the
   already-verified platform dependency, rather than trusting npm rebuild's
   own exit code alone. New tests cover both the matching case and a
   fabricated mismatch (a postinstall that writes fixed wrong content
   instead of resolving the platform dependency).

4. Low: a re-run of install_npm for the same id/version into an
   already-populated final_prefix would install directly into the live
   prefix, leaving a window where the wrapper's freshly updated files point
   at whatever platform dependency npm's own install just auto-fetched,
   unverified, before the fix-up runs. install_npm now stages into a fresh
   directory under stage_dir when a platform_dependency is involved and
   moves the finished, fully verified result into the final prefix only
   after install_platform_dependency succeeds (mv on the same filesystem is
   a single atomic rename), so the live prefix is always either the
   complete old install or the complete new one. New test runs install_npm
   twice for the same id/version and checks both succeed, the staging
   directory never survives either run, and the final content is still
   fully verified after the re-run.

5. Low, launchd-agents.sh remove: a launchctl print failure was treated as
   "not loaded" for ANY nonzero exit; now only 113 ("Could not find
   service") is treated that way. Any other print failure (permission,
   launchd itself unresponsive, ...) keeps ownership and reports, rather
   than guessing either way. The shared _launchctl_shim test helper was
   upgraded to track loaded state per label (bootstrap sets a marker, print
   reads it, bootout clears it) to match this more realistic behavior, and
   every affected test's expected launchctl-print/bootout counts were
   recomputed and verified against the actual, now-more-realistic call
   sequence.

6. Low: pins-macos-arm64.json's top-level `source` field and
   tests/test_adoption_bootstrap_macos.py's
   test_client_npm_pins_name_their_unpinned_darwin_arm64_dependency still
   described/named the round-1 `<alias>@file:<path>` design and the
   "unpinned" framing. Both updated; the test (renamed) now also asserts
   each platform_dependency carries a real sha256, not just a mention in
   install_note.

7. Nit: render_launchd.py's plistlib.dumps call moved inside the same try
   as _substitute, so a control character in a substituted value (which XML
   plist text content cannot represent) raises the same RenderError every
   other failure here does, instead of an uncaught ValueError. New test
   confirms this with a real control character.

Evidence class: the shadow-copy defeat (closed by the coordinator's own
independent Opus re-review, which downloaded and hashed the real tarballs
itself), the NODE_PATH decoy, the resolved_package check and the atomic
staging are all measured against real npm/Node on this host; the launchd
ownership/reinstall fixes remain offline local_integration and structural
evidence pending the next hosted macos-15 run (hosted CI at 522bbe3 is
already green, 15/15, per the coordinator).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Regenerate the new-host grand list after #101/#102 and re-register round-3 evidence

Reflects the macOS pin/socraticode changes in catalogs/landscape/new-host-grand-list.json
and docs/new-host-grand-list.md (scripts/component_matrix.py --write produced no diff;
scripts/new_host_grand_list.py --write did). Also re-registers the six files changed by
round 3 (bootstrap-macos.sh, launchd-agents.sh, pins-macos-arm64.json,
test_adoption_bootstrap_macos.py, test_adoption_launchd.py, render_launchd.py), which had
been left registered at their pre-round-3 hashes after the origin/main merge. Verified via
scripts/component_matrix.py --check, scripts/new_host_grand_list.py --check,
scripts/evidence_manifest.py --check and scripts/validate.py, all passing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Round 3b: canonicalize paths across a symlinked ancestor; launchd install ordering/gating

bootstrap-macos.sh (Opus M1 / hosted macos-15 run 35820422561; Codex round-3
non-atomic-swap Medium):
- Add canonical_path(), a bash 3.2-safe, realpath(1)-free canonicalizer
  (python3 os.path.realpath, falling back to cd -P/pwd -P and a manual
  ancestor walk for a path that does not exist yet). A real Mac's /var is a
  symlink to /private/var, and Node's require.resolve realpath-resolves
  symlinks by default when it locates a module, so a prefix built under a
  symlinked ancestor (TMPDIR-based staging, or a test fixture root) never
  string-equaled the hand-built expected path. install_platform_dependency
  now canonicalizes its own prefix argument on entry and the resolved path
  Node reports; the post-install resolution assertion also realpaths both
  sides via fs.realpathSync. install_npm canonicalizes final_prefix and the
  staged/live prefix. Reproduced off-Mac with two new tests that build the
  fixture prefix/ecosystem_root under a symlinked tmp root (mirroring /var ->
  /private/var); both fail with "function canonical_path not found" (or, pre
  this fix, the exact "resolved to /priva..." mismatch) without the fix.
- L1: replace the non-atomic `rm -rf "$final_prefix"; mv` prefix swap with
  rename-old-aside, move-staged-in, then delete-old -- a crash between the
  two never leaves final_prefix's own bin_dir symlinks pointing at nothing
  with the previous, fully verified install unrecoverably gone.
- N2: rewrite install_platform_dependency's own design comment to state the
  exact-match containment condition and the resolved-package-name check, and
  the new canonicalization; update macos-arm64.md's platform_dependency
  paragraph (Codex), replacing its stale description of a lockfile/
  `_integrity`-based verifier that no longer exists.

launchd-agents.sh (Opus + Codex round-3):
- L2: gate the pre-reinstall launchctl print/bootout probe on
  is_enabled_label, so install never unloads a label it does not own merely
  because it happens to be loaded.
- L3: move directory creation and the owned-destination backup before that
  probe, so a failure in either never leaves an already-unloaded service
  down with nothing done to restore it.
- L4: add wait_until_unloaded, a bounded poll of launchctl print (checked
  before any sleep, so the common case costs no wall-clock time) after a
  successful bootout, since bootout can return before real launchd teardown
  finishes; refuse rather than bootstrap over a service that never actually
  reported unloaded.
- L5: install now treats only launchctl print exit 113 as confidently "not
  loaded", matching remove; any other failure refuses instead of silently
  assuming unloaded.
- L6: add explicit remove coverage for print exiting 5 (keeps ownership and
  the plist; the existing code path was already correct, now asserted).
- N1: move the temporary pre-overwrite backup out of
  ~/Library/LaunchAgents (which launchd auto-loads at login) into
  state_dir/backups, with a trap-based cleanup (pending_backup_plist) as the
  safety net for a premature exit.
- Five new behavior tests plus three structural ones; three of the five
  (L2, L4, L5) fail against the pre-fix script, confirmed by stashing it.

Full acceptance re-run on this commit: named macOS/launchd modules and the
full suite (2688 tests) under bash 5, the same bash-3.2/launchd-gated subset
(50 tests) under a real compiled GNU bash 3.2.0, --plan clean under both bash
versions (shimmed Darwin/arm64), shellcheck and zizmor clean, git diff
--check clean, scripts/validate.py and the catalog/foundation/component-
matrix/grand-list/evidence-manifest checks all passing, and guarded gitleaks
(dir scan and the git-history scan scoped to this branch's own commits since
origin/main) finding no leaks. No macOS host ran any of this.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Round 3c: rollback the failed prefix swap; stop the launchd trap truncating the live plist

bootstrap-macos.sh (Codex round-3 verification of 5310fe5, Medium 1):
- The rename-old-aside/move-staged-in swap had no rollback if the second
  move (staged prefix into final_prefix) itself failed: final_prefix was
  left absent with every bin_dir symlink into it broken, even though the
  previous, fully verified install still existed on disk as
  final_prefix.previous.$$. Failure injection (a shimmed mv failing on any
  "*-staged" source) confirmed this on 87327dc: exit 1, final_prefix
  absent, only the ".previous." copy remaining. Fixed by moving that
  previous copy straight back into final_prefix's place before returning,
  so a failed move loses only the newly staged prefix, never the working
  one bin_dir already pointed at.
- Corrected the design comment's "atomic" claim (also flagged) to say what
  is actually guaranteed: this swap is recoverable, not atomic; a
  genuinely atomic swap would need a versioned directory plus a single
  symlink flip, a larger restructuring this fix does not make.
- New failure-injection test proves the live prefix (and require.resolve's
  own resolution into it) survives a failed final move with the original
  content intact; confirmed failing against the pre-fix script by stashing
  it (reproduces the exact "absent, previous-only" outcome Codex reported).

launchd-agents.sh (Codex round-3 verification of 5310fe5, Medium 2):
- The round-3b N1 trap was itself a regression: `cp -- "$source_plist"
  "$dest_plist"` wrote directly into the live destination, so a copy that
  failed partway (disk full, an interrupted write) could truncate it and
  exit via `set -e` before any of cmd_install's own restore branches ran --
  and the trap only DELETED the pending backup instead of restoring it,
  losing both the live plist's content and its one recovery copy (the
  pre-N1 script at least kept the backup in this scenario). Differential
  injection (a shimmed cp truncating any destination under
  .../LaunchAgents/*.plist) confirmed the regression against 87327dc's
  pre-3c script: destination truncated, backup deleted.
- Fixed two ways: (1) install now copies to a same-directory "*.new.*" temp
  file and renames it into place, so a failed copy never reaches
  dest_plist at all; (2) the EXIT trap now restores pending_backup_dest
  from pending_backup_plist before removing the backup, rather than merely
  deleting it, covering whatever other failure might still reach it.
- New failure-injection test (the same differential shim) proves the fixed
  script's reinstall succeeds cleanly -- the injection never fires at all,
  since the vulnerable direct-to-destination copy no longer exists --  and,
  confirmed by stashing the pre-fix script, that the same injection there
  reproduces the exact truncate-and-delete-the-backup regression.

Full acceptance re-run: named macOS/launchd modules (115 tests) and the
bash-3.2/launchd-gated subset under a real compiled GNU bash 3.2.0 all pass;
bash -n and shellcheck clean on both scripts; scripts/validate.py and the
evidence manifest check pass. No macOS host ran any of this.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Re-register test_workflow_hardening.py after merging origin/main (#108)

scripts/validate.py flagged a hash/byte-count mismatch for
tests/test_workflow_hardening.py (part of #108's GitHub automation
closure) immediately after merging origin/main; re-registered via
host_receipts.register_file and re-normalized with evidence_manifest.py
--write. All other post-merge checks (component_matrix, new_host_grand_list,
evidence_manifest --check, validate_catalogs, validate_foundation,
build_verdicts --check, the full 2796-test suite) pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Round 3d: genuine recovery for both swaps (rollback, was_loaded/reload, atomic flip)

Codex's round-3 verification of 5310fe5 returned pass-with-findings for
everything except the two swaps' own recovery paths, which it reproduced by
direct failure and signal injection. Fixed by design (step tables added as
comments next to each swap; every row tested, including a failed rollback
and SIGTERM mid-swap sent via kill -TERM $PPID from a shim, matching the
coordinator's required method) rather than patched incrementally again:

bootstrap-macos.sh -- replaced the rename-aside-then-move-in swap with the
Homebrew Cellar/opt pattern: install into a versioned, never-reused
directory (tools/<id>-<version>-<stamp>); final_prefix becomes a SYMLINK,
flipped by a single python3 os.replace (genuinely atomic, unlike `ln -sfn`,
which is unlink-then-symlink). bin_dir's own symlinks are unchanged -- they
point into final_prefix/bin/* and transparently follow it through the
extra indirection, so they never need re-creating on a later flip. The one
non-atomic step this still has is a real, pre-3d final_prefix directory's
one-time migration aside (rename() cannot replace a directory with a
symlink in one call); the top-level cleanup() trap restores it if a
failure or a signal lands in that window, reporting the exact manual `mv`
if the restore itself also fails.

Also fixed along the way: final_prefix must never be canonicalized (it
silently resolved straight through the symlink to whatever it currently
targets, defeating the whole design -- caught by the first new failure-
injection test); pending_migration_prefix/_dest must be set BEFORE the
migration mv, not after (a signal in the gap between the mv completing and
the assignment left the trap unable to find what to restore -- caught by
the SIGTERM test); the test harness (_run_install_npm) needed the same
top-level cleanup()/EXIT trap the real script installs, which it had never
included, so failure/signal recovery was silently never exercised at all
until this was added.

launchd-agents.sh -- the EXIT trap now restores only by same-directory
rename (never `cp`, whose own failure was previously ignored before
deleting the backup unconditionally); deletes the backup only after that
rename succeeds, otherwise keeps it and reports the exact manual command;
records was_loaded (via pending_backup_reload_needed, set right after a
confirmed bootout) so restoring the backup from ANY exit path -- not just
the explicit bootstrap-failure branch -- also reloads the service, not
just its file. cmd_install's own explicit restore branch was removed
entirely in favor of the trap (round 3c/3b had two independent
implementations of the same recovery that had independently drifted
different bugs; one is now authoritative). The backup itself moved back
into dest_plist's own directory (reverting round 3b's N1 relocation under
state_dir): `mv` across a filesystem boundary silently falls back to a
non-atomic copy-then-unlink, so the restore-by-rename this fix depends on
is only genuinely atomic when the backup is guaranteed to be on the same
filesystem as dest_plist, which only its own directory can guarantee.

New failure/signal-injection tests (7 total: 3 for the atomic flip, 4 for
launchd's trap), each confirmed to fail against 96f1abe by stashing it.

Full acceptance re-run on this commit: full suite (2801 tests) passes
except the one evidence-registration lag scripts/validate.py itself always
flags immediately after an edit (resolved by the registration commit that
follows); named macOS/launchd modules (120 tests) and the bash-3.2/
launchd-gated subset (54 tests) under a real compiled GNU bash 3.2.0 all
pass; --plan clean under both bash versions; shellcheck and zizmor clean;
git diff --check clean. No macOS host ran any of this.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Round 3e: idempotent convergence for launchd, ownership-checked pruning for bootstrap

Codex's round-3e review of c18aea3 found two new Highs and four Mediums --
each earlier round's hand-written rollback had closed one edge case and
opened another (an unowned rm -rf, a partial hard-copy backup, a reload
skipped at a step boundary, cleanup left armed after success, a signal
racing a foreground mv, an unguarded rm aborting recovery under set -e).
Per the coordinator's explicit direction, this round changes the approach
instead of patching further: simplify toward idempotent convergence, the
way `brew services` works, with step tables and signal-injection tests
proving every row.

Both scripts now explicitly trap INT, TERM and HUP (`exit 130/143/129`
only; recovery logic stays in the EXIT handler alone) -- bash only
reliably defers a caught signal to a completed step when the signal is
explicitly trapped, not left at its default disposition (confirmed
directly: an untrapped SIGINT sent to a bash script with only an EXIT trap
registered does not interrupt it at all in this environment, while an
untrapped SIGTERM does -- the asymmetry Codex's finding 5 was pointing at).

bootstrap-macos.sh (High #1, unowned rm -rf): new prune_old_version()
deletes a superseded versioned directory only when its canonical form is a
direct child of the canonical tools/ directory with a name matching
<id>-<version>-<stamp>; an external target, a relative "../" target, a
non-matching name, or a symlink loop (canonical_path's python3
os.path.realpath handles cycles without hanging) is logged and left in
place, never deleted. Pruning is now genuinely best-effort: a rejected or
failed deletion warns and never fails the install. Four new tests, one per
adversarial target shape, confirmed to fail against c18aea3 (the function
does not exist there).

launchd-agents.sh -- full redesign, net roughly flat in size (+22 lines)
despite ~390 changed, since one pure function replaces four tracking
markers (pending_backup_plist/_dest/_reload_needed, pending_dest_tmp) and
their two independent, historically drifting recovery implementations:
- backup_plist is a hard link (`ln`, High #2), not `cp`: it either exists
  completely or not at all, so a partial copy can never overwrite the
  intact original. Same directory as dest_plist (round 3d), never this
  script's own state_dir.
- Staging (cp to dest_plist.new, then mv over dest_plist) is unchanged in
  shape but the name lost its PID suffix, matching the backup's own
  naming, so a stale artifact from a crashed run is visible to and
  reconciled by the very next run, not just the same process.
- reconcile_install replaces every marker: it reads dest_plist's identity
  against backup_plist (-ef, since the backup is a hard link) and
  launchd's own current loaded state, and converges to whichever ONE
  action that state calls for (drop the backup, reload, restore-and-
  reload, or -- if nothing converges -- keep the backup and name the exact
  re-run command). It is the ONLY recovery logic: called explicitly after
  a normal attempt, and identically as the EXIT trap for anything that
  interrupts one. `set +e` inside it, explicit checks throughout, per the
  coordinator's instruction that the handler must never itself become
  fatal under the script's own `set -e`.
- record_enabled_label now runs as soon as install actually commits to a
  label (not only on success), so a label that ends up needing attention
  is never re-refused as unowned on retry; a was_already_enabled snapshot,
  taken before that point, keeps the L2 bootout gate correctly checking
  ownership as it stood BEFORE this run, not after.

New tests: a shared _signal_shim helper plus a parameterized step x signal
matrix (ln and mv, TERM and INT) proving retry-converges for launchd; a
dedicated SIGTERM/SIGINT-parameterized migration test and three more for
bootstrap-macos.sh's prune_old_version. Every new/changed test confirmed
against c18aea3 by stashing the fix; several rewritten round-3b/c/d tests
(now asserting the new design's own honest behavior -- e.g. a fresh
install that never loads is left on disk with a manual-retry message,
never silently deleted) checked to still exercise the coordinator's
original findings under the new mechanism.

Full acceptance re-run on this commit: full suite (2821 tests) passes
except the one evidence-registration lag scripts/validate.py itself always
flags immediately after an edit (resolved by the registration commit that
follows); named macOS/launchd modules (126 tests) and the bash-3.2/
launchd-gated subset (56 tests) under a real compiled GNU bash 3.2.0 all
pass; --plan clean under both bash versions; shellcheck and zizmor clean;
git diff --check clean; guarded gitleaks clean. origin/main had not moved
since it was already merged into c18aea3's own parent, so no new merge was
needed this round. No macOS host ran any of this.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Round 3f: abort preflight on unresolved reconcile, path-verify launchctl state, poll through delayed teardown

Codex verified d51f062 FAIL with a High and two Medium regressions in
launchd-agents.sh's reconcile_install/cmd_install, plus an open delayed-
teardown boundary; bootstrap-macos.sh itself was confirmed sound and is
untouched here (its test matrix gained the same row/signal coverage).

- High: reconcile_install's own preflight failure is no longer ignored
  (`|| exit 1`, not `|| true`): a retry that still cannot converge now
  aborts before ever reaching the backup-overwriting `ln`, so the last
  known-good backup can no longer be silently destroyed and re-linked to
  unusable content.
- Medium (tri-state print): launchctl_label_state() reports loaded_here,
  loaded_elsewhere, not_found or unknown instead of a binary loaded check;
  unknown now takes no destructive action anywhere and keeps the backup.
- Medium (unrelated loaded label): ownership is recorded only inside
  reconcile_install, once a path match confirms dest_plist itself is what
  launchd has loaded -- never earlier -- and both reconcile_install's
  fresh-install path and cmd_install's own bootout gate refuse outright on
  loaded_elsewhere, on the first invocation and every retry.
- Delayed teardown: reconcile_install re-polls (reusing wait_until_unloaded)
  when this run's own bootout is still pending before trusting a "loaded"
  read, instead of declaring convergence with zero reload attempts.

Added four named regression tests reproducing Codex's three exact scenarios
plus the delayed-teardown case, each verified via differential (git stash)
to fail on d51f062 and pass here. Expanded the checked-in matrix to cover
every row of both step tables under failure, SIGTERM and SIGINT (was two
launchd rows and one bootstrap row), reusing and extending the existing
signal-shim helpers rather than one bespoke script per case.

Updated the in-script step table for the ownership-timing and preflight-
abort changes, and recorded the delayed-teardown re-poll as an untested
boundary for real-Mac acceptance in the platform notes: modelled against a
mock launchctl and real bash signal injection, never observed on native
launchd. Verified with the full adoption test suite (61 launchd + 78
bootstrap-macos tests) under both bash 5.2 and a real compiled bash 3.2.0,
shellcheck, zizmor, a manual bash-3.2 offline render/lint/install/status/
remove cycle, a manual bash-3.2 real-SIGTERM injection during install,
scripts/validate.py and friends, and guarded gitleaks (dir + git log range)
-- all clean. origin/main has not moved past this branch's base; no merge
needed. Evidence class unchanged: Linux-side local_integration/structural
plus real npm/Node/bash-3.2/signal behavior measured directly on this host;
no macOS host has run any of this.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Scope the pre-reinstall load-check test to cmd_install's own body

test_install_gates_its_pre_reinstall_load_check_on_ownership matched
cmd_remove's textually similar is_enabled_label + launchctl print gate
instead of cmd_install's round-3f one, so it passed for the wrong reason.

Added _shell_functions (the same top-level-function extraction helper
tests/test_adoption_bootstrap_macos.py already uses) to scope the
assertion to cmd_install's own body, and rewrote it to check the round-3f
gate directly: launchctl_label_state is called before any bootout, and the
loaded_elsewhere and unknown case arms both refuse (exit 1, no bootout)
rather than proceed.

Verified the new test actually exercises cmd_install and not cmd_remove by
temporarily replacing cmd_install's gate with the old pre-round-3f binary
launchctl-print check (no path verification, no loaded_elsewhere/unknown
refusal) and confirming the test fails against that mutation, then
restoring the gate byte-for-byte (diff against HEAD is empty for
launchd-agents.sh; only the test file and evidence manifest changed here).

Re-ran the full test_adoption_launchd suite (61/61) under bash 5.2 and a
real compiled bash 3.2.0, plus bash -n on both interpreters, shellcheck,
scripts/validate.py, evidence_manifest.py --check (re-registered via
host_receipts.register_file) and guarded gitleaks dir -- all clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Round 3g: replace launchd's transactional backup/reconcile with brew-services semantics

Codex's review of 0e1e0d4 found four new Medium regressions in the round
3b-3f backup/reconcile design (remove booting out an unrelated service on
stale state; a print error on fresh install deleting an already-loaded
plist; a signal during bootout or restore leaving the service stopped with
no retry built in; an ownership-write failure falsely reporting success) --
the fifth straight round to find a new recovery defect in that machinery.
Rather than patch a sixth, launchd-agents.sh now converges on the
maintained upstream pattern: Homebrew/brew Library/Homebrew/services/
cli.rb @ 8e3a5dc0a7 (2026-09-07), which has no backup, no rollback and no
ownership file, and recovers by re-running the command.

- Removed record_enabled_label, is_enabled_label, forget_enabled_label,
  the enabled-labels state file, reconcile_install, was_already_enabled
  and the hard-linked backup entirely.
- Stateless ownership: "ours" means the label is one of this script's own
  AND, if it is loaded at all, launchctl print's own "path =" line names
  this script's destination plist -- never a stored record. install and
  remove share the exact same launchctl_label_state check (tri-state plus
  elsewhere, unchanged from round 3f) and both refuse outright, touching
  nothing, on loaded_elsewhere or unknown.
- install: stage a temp file beside the destination and lint it; if
  loaded_here, bootout and wait until confirmed unloaded (refusing
  otherwise); rename into place; launchctl enable + bootstrap; on
  bootstrap failure, leave the new plist in place (no rollback -- there is
  nothing to roll back to) and print the exact recovery commands.
- remove: on loaded_here, bootout, wait, then delete; on not_found, delete
  directly if a file is present; refuse on loaded_elsewhere or unknown.
- The EXIT trap now only removes the current run's own still-inert temp
  file; INT/TERM/HUP stay explicitly trapped so bash always defers to a
  completed step.

Replaced every round 3b-3f backup/reconcile test with a convergence matrix
(cp/bootout/mv/bootstrap x {failure, TERM, INT}, then a fault-free install
retry converging to NEW loaded, then a fault-free remove leaving nothing
loaded and no file) plus named tests for loaded_elsewhere refused by both
install and remove, unknown touching nothing on both, and Codex's exact
remove reproduction (a previously installed label's plist on disk, its
label now loaded from a different, unrelated path). All new/changed tests
verified via differential (git stash on the script alone) to fail on
0e1e0d4 (19 of 55 tests fail there, the rest being unaffected render/
template/structural checks); all 55 pass here.

adoption/platforms/macos-arm64.md carries the dated decision (chosen: this
design, with the Homebrew citation; rejected: the transactional design,
because rounds 3b-3f each found a new recovery defect in it; would
overturn: a real-Mac observation where a re-run does not converge, or a
requirement to preserve a locally edited plist) and keeps the untested
delayed-teardown-polling boundary note, updated for the new call sites.

launchd-agents.sh shrank from 758 to 593 lines (38310 to 28554 bytes)
against 0e1e0d4, removing the transactional machinery outright rather than
adding another layer of recovery logic on top of it.

Merged origin/main (bb5228d, unrelated: release-tag pin bump) -- no
conflicts with this file set. Verified: full test_adoption_launchd (55/55)
and test_adoption_bootstrap_macos (78/78, unaffected) under bash 5.2.21
and a real compiled bash 3.2.0; a manual bash-3.2 offline render/install/
status/remove cycle and a manual bash-3.2 real-SIGTERM injection during
the rename step; shellcheck, bash -n on both interpreters, zizmor,
git diff --check, scripts/validate.py and friends (evidence re-registered
via host_receipts.register_file), and guarded gitleaks (dir) -- all clean.
Evidence class unchanged: Linux-side local_integration/structural plus
real bash-3.2/signal behavior measured directly on this host; no macOS
host has run any of this.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Round 3h: pin and download the llama-embed model, add the embedding acceptance check

A 2026-09-23 readiness audit found two defects in PR #94:

1. The llama-embed launchd agent's template ran llama-server with no
   model argument at all -- KeepAlive would restart it in a loop forever
   on a real Mac, never actually serving embeddings. Fixed:
   - com.native-stack.llama-embed.plist.template now passes -m
     ${EMBED_MODEL_PATH}.
   - adoption/pins-macos-arm64.json gained a separate top-level "models"
     section (never "tools", so it is never swept into
     adoption/manifest.json's protected profile-component coverage) for
     ggml-org/embeddinggemma-300M-GGUF @ 0f741b5a6585bd53aeb15cd1372c56f2a0f65e12,
     embeddinggemma-300M-Q8_0.gguf, sha256-pinned, with its own
     huggingface_lfs_oid_plus_local_rehash checksum source.
   - adoption/bootstrap-macos.sh gained install_embed_model, a dedicated
     download-and-verify step (reusing fetch(), the same helper every
     other pin already uses) run unconditionally after the tools[]
     loop, fail-closed on a null or mismatched sha256, into
     $ECO_INSTALL_ROOT/state/models/.
   - launchd-agents.sh's cmd_render and adoption/hosts/macos-example.json
     both carry EMBED_MODEL_PATH now, alongside the existing HOME/
     ECO_ROOT/AI_MEMORY_URL host values.
   - New tests: the rendered plist's -m argument matches EMBED_MODEL_PATH
     (test_adoption_launchd.py); install_embed_model, extracted verbatim
     and run against a real (curl-shimmed) fetch(), installs on a
     matching digest and fails closed on a mismatch or a null sha256
     (test_adoption_bootstrap_macos.py); the models[] pin's own shape,
     URL-matches-revision and separation from tools[].

2. No Linux reference vector existed to compare a macOS response
   against. Fixed:
   - evidence/artifacts/macos-embed-reference-20260923/ carries the
     captured reference (llama.cpp b11057 ubuntu-x64 CPU build, the
     identical pinned GGUF, 768 dimensions, L2-normalized, repeat cosine
     1.0, full request/response provenance), registered via
     host_receipts.register_file.
   - tools/adoption/embed_acceptance.py (stdlib only) sends the
     reference's exact request body to a running llama-server, checks
     dimension and cosine >= threshold against the reference embedding,
     and prints a JSON result; understands both the OpenAI-compatible
     and llama.cpp-native response shapes, and reports a network or
     malformed-response failure as JSON, never a bare traceback. Six new
     tests (a local stdlib HTTP server stands in for llama-server) cover
     an exact match, both response shapes, a dimension mismatch, a
     below-threshold cosine, and a connection failure.
   - .github/workflows/adoption-bootstrap.yml's bootstrap-macos job now
     caches the pinned model (actions/cache, keyed on its sha256, not a
     version tag), installs and health-checks the llama-embed launchd
     agent, runs embed_acceptance.py against it, uploads the JSON result
     and the startup log, and best-effort-greps the log for which
     backend it reports -- labelled explicitly as hosted-runner evidence
     only (a GitHub-hosted macOS runner is CPU-only, never a claim of
     Metal actually engaging; a real Mac run is what would show that).
   - adoption/platforms/macos-arm64.md's "What a hosted run proves" and
     "Acceptance test for the 24 GB default" sections both updated: the
     reference vector now exists, and the one command to run it is
     documented, still bounded to "unrun on any Mac, hosted or
     otherwise" for the acceptance itself.

Verified: full test_adoption_bootstrap_macos (93/93) and
test_adoption_launchd (56/56) under bash 5.2.21 and a real compiled bash
3.2.0; bash -n and shellcheck clean on both scripts; zizmor clean on the
workflow; git diff --check; scripts/validate.py and friends (evidence
re-registered via host_receipts.register_file); guarded gitleaks dir --
all clean. Evidence class unchanged and explicit throughout: Linux-side
local_integration/structural, a real captured Linux reference vector, and
(once this round's hosted CI step runs) CPU-only hosted-runner evidence;
no macOS workstation has run any of this.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Round 3i: canonical-path comparison, EINPROGRESS handling, cache-before-verify, embed_acceptance transport errors, macOS recording-tooling CI

Codex's review of e9334e9 returned FAIL with three Medium findings and one
Low, all with narrow, well-isolated fixes:

1. launchd-agents.sh's loaded-state classifier used `-ef` (device+inode
   identity), so a DISTINCT hard-linked filename sharing dest_plist's own
   inode incorrectly compared loaded_here, letting remove bootout and
   delete a plist for a label actually loaded from an unrelated,
   hard-linked path. Replaced with canonical_plist_path: resolves symlinks
   in a path's parent directory only (cd -P, falling back to python3's
   os.path.realpath), then reattaches the final component unresolved --
   neither mechanism ever collapses a hard link's own distinct pathname
   the way inode identity did, since a hard link has no stored "canonical
   name" to resolve to.

2. A `launchctl bootout` returning EINPROGRESS (teardown already under
   way, from this run's own retry of an earlier interrupted attempt, or
   another process) was treated as an unconditional failure, abandoning
   an in-progress teardown with no wait and no reload ever attempted. The
   pinned Homebrew cli.rb @ 8e3a5dc0a7 (services/cli.rb, ~line 311)
   retries on exactly Errno::EINPROGRESS (Darwin 36) for the identical
   reason. New shared bootout_and_wait absorbs it the same way, then
   polls (reusing wait_until_unloaded) rather than refusing outright.
   Self-discovered along the way: the first implementation captured its
   result via `result="$(bootout_and_wait ...)"`, which forks a subshell
   -- empirically verified that a signal delivered while a foreign
   command runs INSIDE that subshell is silently swallowed there and
   never reaches this script's own INT/TERM/HUP traps at all. Fixed by
   using a plain function call plus a global result variable instead,
   keeping bootout_and_wait in the same process as its caller.

3. adoption-bootstrap.yml's embedding-model cache restore ran AFTER
   bootstrap's own checksum verification and used a literal hardcoded
   sha256 as its key. A restore running that late could silently
   overwrite the just-verified model with stale cached bytes, with
   nothing left to re-check them, and a changed pin would keep matching
   the same stale cache entry. Fixed: a new step reads the pin's sha256
   into a step output, the cache restore (now keyed on that output) moves
   to before the bootstrap step, and bootstrap-macos.sh's own fetch()
   (unchanged, already correct) transparently re-downloads and
   re-verifies any mismatched destination on the exact same path a
   genuine cache miss takes.

4. embed_acceptance.py caught urllib.error.URLError around the request,
   but response.read() can raise a bare TimeoutError (a stall mid-body,
   past urlopen's own connect-phase timeout) or http.client.
   IncompleteRead (the server closing before its declared Content-Length
   was satisfied) -- neither caught, leaving stdout empty and the process
   exiting via an uncaught traceback instead of the documented JSON
   failure contract. Both now caught explicitly, plus a catch-all OSError.

Also, this round (peer-update-audit gap adoption_macos, medium): the
catalog's own recording and verdict scripts (host_receipts.py,
component_matrix.py, new_host_grand_list.py, build_verdicts.py,
validate_convergence.py, release_due.py) had never run on macOS CI or
against macOS's own system Python -- every existing gating job for them
runs on ubuntu-24.04 only. adoption-bootstrap.yml's validate-macos job now
gates on all six, then runs a recording smoke: installs the profile,
records a real host_receipts.py record receipt (component codex,
--from-stack-commands, --evidence-class native_proven -- hosted-runner
evidence only, never a workstation acceptance receipt) inside a throwaway
copy of the checkout ($RUNNER_TEMP/rec, never the real one, so nothing is
committed), then re-validates and re-checks the matrix/grand-list in that
same copy. Runs once against the manifest-pinned Python line and once
against the runner's own system /usr/bin/python3 if it meets the newly
declared Python 3.9 floor (adoption/bootstrap.md, adoption/platforms/
macos-arm64.md's new "Recording and verdict scripts" section), skipping
with a message otherwise.

Added named tests for all five changes (a hard-linked alias never
compares loaded_here; a retry converges while an earlier run's teardown
is still asynchronously in progress, fixing the bootout-signal test shim
Codex flagged for clearing the marker before signalling and so hiding the
async case; CI step order and the derived cache key; two embed_acceptance
transport-failure modes over a raw-socket server; structural checks on
the new validate-macos steps), each verified via differential (git stash)
to fail on e9334e9 or the pre-fix file. Full suites: test_adoption_launchd
59/59, test_adoption_bootstrap_macos 107/107.

Verified: both scripts under bash 5.2.21 and a real compiled bash 3.2.0
(syntax and shellcheck), PLUS two direct, real-SIGTERM/real-shim
reproductions under bash 3.2 outside the Python harness (the hard-link
refusal, and the EINPROGRESS retry actually converging); zizmor and
actionlint both clean on the workflow; git diff --check; scripts/
validate.py and friends including host_receipts.py validate and
build_verdicts.py --check (evidence re-registered via host_receipts.
register_file); guarded gitleaks dir -- all clean. origin/main has not
moved past this branch's base; no merge needed. Evidence class unchanged:
Linux-side local_integration/structural plus real bash-3.2/signal
behavior measured directly on this host; the new CI steps are themselves
unrun until a real macos-15 job executes them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Fetch full history in validate-macos so host receipt catalog revisions resolve

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Resolve symlinked plist leaves in canonical_plist_path; make the incomplete-body test drain the request

canonical_plist_path now follows a symlinked leaf (bounded at 40 hops) so a
plist launchctl reports by its resolved target compares loaded_here, while a
hard link stays a distinct path (Codex round-3i finding 1). The raw-socket test
server drains the whole request and half-closes, so the client sees
IncompleteRead rather than an occasional ECONNRESET (30/30 runs pass).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Skip the PyYAML structural workflow tests where PyYAML is absent (repository convention)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Give the macOS recording smoke a per-run --identity (required since #117)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Record the 2026-09-23 hosted macOS smoke receipt (launchd, embedding acceptance, recording smoke)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Round 3j: fix 8 unresolved Codex-connector review threads blocking the #94 merge

1. launchd-agents.sh:77 -- llama-embed was left out of the default install/
   status/remove label set, even though bootstrap now downloads and
   sha256-verifies its model unconditionally (round 3h). Added it to
   default_labels; --label was the only way to reach it before this.

2. launchd-agents.sh:526 -- install created the plist's own declared
   StandardOutPath/StandardErrorPath/WorkingDirectory BEFORE the ownership
   check, so a refused install (loaded_elsewhere or unknown, which must
   touch nothing) still left new, empty directories behind. Moved
   directory creation to after the ownership check's case statement.

3. macos-arm64.md:217 -- "What a hosted run proves" was stale (cited a
   2026-09-22 run, claimed launchd and the embedding acceptance were both
   still unrun). Rewritten for run 35875188590 at head 75a6e0d: launchd
   genuinely bootstrapped and booted out both qdrant and llama-embed;
   qdrant health reported 1.19.1; the embedding acceptance passed with
   cosine 0.99944 and the startup log reports Metal device MTL0 engaging
   for the layers it covers; the recording smoke passed on both
   actions/setup-python 3.13.15 and the runner's own system
   /usr/bin/python3 3.9.6. Kept the hosted-runner-only scope throughout --
   no platform_status change, no workstation acceptance claim.

4. qdrant.plist.template:28 (the most important one) -- nothing on a
   clean host ever created $ECO_ROOT/config/qdrant.yaml; only the hosted
   CI step's own disposable file did. Added provision_qdrant_config to
   bootstrap-macos.sh: mirrors examples/qdrant.yaml.example exactly
   (storage under $ECO_ROOT/state/qdrant, loopback, port 16333 -- the
   same port adoption/hosts/macos-example.json's QDRANT_URL and
   recipes/README.md both name), plan-mode aware, and never overwrites an
   existing file. The hosted launchd step no longer writes its own
   disposable config or uses port 6333; it uses the provisioned one at
   port 16333, exactly what a real install would use.

5. ai-memory.plist.template:30 -- added --workspace local --project
   native-agent-stack, matching the selected native recipe's own serve
   invocation byte for byte (recipes/README.md:244; the same scope
   docs/foundation-stack.md:51-52 and every other documented ai-memory
   invocation in this project already use).

6. adoption-bootstrap.yml:13 -- the paths triggers covered adoption/**
   and tools/adoption/** but missed most of what validate-macos's
   recording-tooling gate and recording smoke actually consume. Added the
   embed reference artifact, host_receipts.py, component_matrix.py,
   new_host_grand_list.py, build_verdicts.py, platform_status.py,
   validate_convergence.py, release_due.py and manifests/evidence.json to
   both push and pull_request path lists.

7. launchd-agents.sh:517 -- lint only checked plist syntax, never that a
   rendered plist's own Label key matches its filename (also the label
   every other subcommand addresses it by, and what launchd itself
   expects to agree). Checked once syntax lint has already passed, via
   pure bash parameter expansion (no new basename/dirname dependency).

8. launchd-agents.sh:620 -- status always exited 0, even when every
   requested label's launchctl print failed (absent, 113, or launchd
   itself unavailable, any other nonzero exit) -- a caller checking only
   the exit code never saw either case. Now returns non-zero whenever any
   requested label's print fails…
@seathatflowsinourveins
seathatflowsinourveins deleted the codex/evidence-convergence-practice branch September 25, 2026 18:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant