Repository navigation
Publish evidence-led convergence practice and retrieval comparisons - #2
Merged
seathatflowsinourveins merged 2 commits intoSep 20, 2026
Merged
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
seathatflowsinourveins
changed the base branch from
codex/mac-worker-timeout-current
to
main
September 20, 2026 02:25
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 23, 2026
…ng for bootstrap Codex's round-3e review of c18aea3 found two new Highs and four Mediums -- each earlier round's hand-written rollback had closed one edge case and opened another (an unowned rm -rf, a partial hard-copy backup, a reload skipped at a step boundary, cleanup left armed after success, a signal racing a foreground mv, an unguarded rm aborting recovery under set -e). Per the coordinator's explicit direction, this round changes the approach instead of patching further: simplify toward idempotent convergence, the way `brew services` works, with step tables and signal-injection tests proving every row. Both scripts now explicitly trap INT, TERM and HUP (`exit 130/143/129` only; recovery logic stays in the EXIT handler alone) -- bash only reliably defers a caught signal to a completed step when the signal is explicitly trapped, not left at its default disposition (confirmed directly: an untrapped SIGINT sent to a bash script with only an EXIT trap registered does not interrupt it at all in this environment, while an untrapped SIGTERM does -- the asymmetry Codex's finding 5 was pointing at). bootstrap-macos.sh (High #1, unowned rm -rf): new prune_old_version() deletes a superseded versioned directory only when its canonical form is a direct child of the canonical tools/ directory with a name matching <id>-<version>-<stamp>; an external target, a relative "../" target, a non-matching name, or a symlink loop (canonical_path's python3 os.path.realpath handles cycles without hanging) is logged and left in place, never deleted. Pruning is now genuinely best-effort: a rejected or failed deletion warns and never fails the install. Four new tests, one per adversarial target shape, confirmed to fail against c18aea3 (the function does not exist there). launchd-agents.sh -- full redesign, net roughly flat in size (+22 lines) despite ~390 changed, since one pure function replaces four tracking markers (pending_backup_plist/_dest/_reload_needed, pending_dest_tmp) and their two independent, historically drifting recovery implementations: - backup_plist is a hard link (`ln`, High #2), not `cp`: it either exists completely or not at all, so a partial copy can never overwrite the intact original. Same directory as dest_plist (round 3d), never this script's own state_dir. - Staging (cp to dest_plist.new, then mv over dest_plist) is unchanged in shape but the name lost its PID suffix, matching the backup's own naming, so a stale artifact from a crashed run is visible to and reconciled by the very next run, not just the same process. - reconcile_install replaces every marker: it reads dest_plist's identity against backup_plist (-ef, since the backup is a hard link) and launchd's own current loaded state, and converges to whichever ONE action that state calls for (drop the backup, reload, restore-and- reload, or -- if nothing converges -- keep the backup and name the exact re-run command). It is the ONLY recovery logic: called explicitly after a normal attempt, and identically as the EXIT trap for anything that interrupts one. `set +e` inside it, explicit checks throughout, per the coordinator's instruction that the handler must never itself become fatal under the script's own `set -e`. - record_enabled_label now runs as soon as install actually commits to a label (not only on success), so a label that ends up needing attention is never re-refused as unowned on retry; a was_already_enabled snapshot, taken before that point, keeps the L2 bootout gate correctly checking ownership as it stood BEFORE this run, not after. New tests: a shared _signal_shim helper plus a parameterized step x signal matrix (ln and mv, TERM and INT) proving retry-converges for launchd; a dedicated SIGTERM/SIGINT-parameterized migration test and three more for bootstrap-macos.sh's prune_old_version. Every new/changed test confirmed against c18aea3 by stashing the fix; several rewritten round-3b/c/d tests (now asserting the new design's own honest behavior -- e.g. a fresh install that never loads is left on disk with a manual-retry message, never silently deleted) checked to still exercise the coordinator's original findings under the new mechanism. Full acceptance re-run on this commit: full suite (2821 tests) passes except the one evidence-registration lag scripts/validate.py itself always flags immediately after an edit (resolved by the registration commit that follows); named macOS/launchd modules (126 tests) and the bash-3.2/ launchd-gated subset (56 tests) under a real compiled GNU bash 3.2.0 all pass; --plan clean under both bash versions; shellcheck and zizmor clean; git diff --check clean; guarded gitleaks clean. origin/main had not moved since it was already merged into c18aea3's own parent, so no new merge was needed this round. No macOS host ran any of this. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 23, 2026
…hd agents (brew-services semantics), embedding acceptance, recording-tooling CI (#94) * Pin socraticode and platform-dependency integrity, add launchd agents for the macOS adoption profile adoption/bootstrap-macos.sh: replace the jq-only Homebrew block with a bash-3.2-safe loop over the six floating formulae (jq, python@3.13, ripgrep, coreutils, restic, shellcheck), installing only what `brew list --versions` reports missing, gated on --skip-system-packages/--plan, and echoed in --plan. socraticode now has a reviewed npm pin (--ignore-scripts, matching the Linux recipe) instead of being a documented, unpinned skip, so documented_unpinned_ids is now empty (guarded for bash 3.2). Add verify_platform_dependency: a fail-closed check, run right after `npm install`, that the darwin-arm64 platform_dependency npm resolves for codex and claude-code (name/version/integrity) matches the host's own .package-lock.json or package.json _integrity. adoption/pins-macos-arm64.json: add the socraticode pin (sha256 verified by downloading the tarball and cross-checking against `npm view socraticode@1.14.0 dist.integrity`) and a platform_dependency entry for each of @openai/codex and @anthropic-ai/claude-code, with npm dist.integrity read 2026-09-23. @openai/codex-darwin-arm64 is not itself a published package name; `npm view @openai/codex@0.155.1 optionalDependencies` shows it as an npm: alias to @openai/codex@0.155.1-darwin-arm64, recorded as such. New adoption/launchd/: three plist templates (qdrant, ai-memory, the llama.cpp Metal embedding server on port 8232) using the same ${NAME} placeholder convention as adoption/templates/, plus launchd-agents.sh (render/lint/install/status/remove; plutil -lint with a plistlib fallback; launchctl bootstrap/print/bootout; remove only boots out a label this script's own state file recorded as enabled, never deleting data). New tools/adoption/render_launchd.py reuses render_config.py's render_one/ load_host_values/parse_set_values. New adoption/hosts/macos-example.json supplies the render_config placeholder schema with /Users/example paths and cites hardware-profiles.json's macos-arm64-48gb-projected id; EMBED_URL is the bare host:port form (not the http://-prefixed form first drafted) since project.codex.config.template.toml already prepends http:// for it. .github/workflows/adoption-bootstrap.yml: add tools/adoption/** to the path filters; add a launchd render/lint/bootstrap-qdrant/health-check/bootout step to the existing bootstrap-macos job; add bootstrap-macos-brew (runs the script without --skip-system-packages, asserting python3.13/rg/restic/ shellcheck/gsha256sum) and validate-macos (the repository's own validators, the adoption test modules as a gate, and one full python3 -m unittest recorded as a non-gating artifact), both on macos-15. adoption/platforms/macos-arm64.md: prerequisites, exit-code table, darwin-arm64 pin table and launchd section updated to match; the "unpinned"/documented-skip wording for socraticode is removed since it is now pinned. tests/test_adoption_bootstrap_macos.py: update DOCUMENTED_SKIPS, the brew block assertions, and add PlatformDependencyVerificationTests (offline, fixture pins + fake .package-lock.json/npm shim; fixture integrity values are assembled from concatenated parts, not written as single literals). New tests/test_adoption_launchd.py covers template rendering/schema, render_launchd.py, launchd-agents.sh structure and an offline render/lint/install/status/remove cycle against a fake launchctl. Evidence class: no macOS host ran anything here. Everything is either local integration (Linux-side unit tests, shellcheck, plistlib parsing) or a hosted macos-15 runner the coordinator dispatches later. adoption/platforms/macos-arm64.md stays drafted_not_accepted, and adoption/manifest.json is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Re-hash the evidence registry and rebuild the ecosystem explorer Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Fix round: real npm/lockfile evidence, bash 3.2 consumption bug, launchd ownership, XML escaping, harden-runner Addresses the Opus evidence review (fail), the Codex cross-family review (fail), and hosted macos-15 CI run 35812470345 on PR #94, which reproduced three of these findings directly and one more (git_revision) on its own. 1. npm platform-dependency verification never could have passed: measured directly against npm 11.19.0 on this project's own host (existing ecosystem prefixes, e.g. mcporter-0.13.13, and a fresh `npm install --global --prefix <dir> is-odd`), a --global --prefix install writes no lockfile anywhere under the prefix, no package.json _integrity, and places a platform-specific optional dependency NESTED under the parent (lib/node_modules/<parent>/node_modules/<platform-dep>), not at its own top-level name -- reproduced here with two real npm-packed local packages (tests.RealNpmLockfileEvidenceTests), not asserted from memory. verify_platform_dependency (post-hoc lockfile/package.json read) is replaced by install_platform_dependency: fetch the platform tarball itself (dist.tarball, added to each platform_dependency pin alongside a freshly re-hashed sha256), verify it with the same fetch() every other pin uses, then `npm install --global --prefix <dir> <alias>@file:<path>` the already-verified local archive -- confirmed with real npm (PlatformDependencyInstallTests) that this places the verified bytes at the alias name regardless of the tarball's own internal package.json name, which is what Node's require() actually resolves by. 2. bootstrap-macos.sh:192's `for allowed_id in "${allowed_unpinned_ids[@]}"` was unguarded: the array's assignment was bash-3.2-safe but this later consumption of it was not, and it aborts with "allowed_unpinned_ids[@]: unbound variable" whenever it is genuinely empty -- exactly the hosted CI failure. Fixed with the same ${arr[@]+"${arr[@]}"} guard used elsewhere. This fix round compiled a real bash 3.2.0 from source (GNU bash source, ftp.gnu.org) to reproduce the abort and confirm the fix dynamically -- not just structurally, which is all bash 5's relaxed nounset handling of empty arrays can ever show. ScriptBehaviorUnderRealBash32Tests runs the exact --plan and exit-3 paths under that real 3.2 binary when BASH32_BINARY is set, or on a real Mac (/bin/bash genuinely is 3.2 there). 2b. tests/test_adoption_launchd.py's end-to-end cycle test asserted the plistlib fallback text unconditionally, which fails wherever plutil is present (a real Mac; not this dev host). Relaxed to a generic success check, with two new dedicated, PATH-controlled tests exercising each path deliberately (test_lint_falls_back_to_plistlib_when_plutil_is_absent, test_lint_prefers_plutil_when_present), so the behavior is proven either way rather than incidentally on whichever host happens to run it. 2c. scripts/adoption_status.py's git_revision(root) compared git's own resolved --show-toplevel path against an unresolved caller root; macOS routes its default tempdir through /var -> /private/var (reproduced here with a plain symlink on Linux), so a direct caller with an unresolved root -- exactly what the existing unit test does -- got None instead of the real SHA on a real Mac. inspect_adoption already resolves root first, so this was latent; git_revision now resolves its own root parameter, closing the gap for every caller. 3. launchd-agents.sh install now refuses to overwrite an existing ~/Library/LaunchAgents/<label>.plist that this script did not itself record as enabled, and creates state/logs and state/qdrant (qdrant's new WorkingDirectory) before installing -- launchd does not create missing parent directories for StandardOutPath/StandardErrorPath. remove keeps ownership (does not forget the label) when launchctl bootout fails, and on success now deletes the copied plist under ~/Library/LaunchAgents (not just booting it out) so RunAtLoad cannot reload it at the next login; it still never touches the component's own data or logs. llama-embed is left out of the default install/status/remove set until a model argument exists (render/lint still cover it, and an explicit --label still installs it). 4. tools/adoption/render_launchd.py rendered plists via a raw string.Template substitution on template text, which is not XML-escaping: a substituted value containing "&" (a real path such as /Volumes/R&D/eco) would corrupt the document. Rendering now parses the template as a plist first (plistlib tolerates the still-literal ${NAME} placeholders as ordinary string content), substitutes inside every string leaf of the resulting structure, and re-serializes with plistlib.dumps, so plistlib's own escaping covers every value; proven with a real "&"-containing host value round-tripped through valid XML. 5. The bootstrap-macos-brew and validate-macos jobs, and the existing bootstrap-macos job (which had none), now start with harden-runner in audit mode: the pinned version's own README (read at its exact pinned commit) states macOS runners get full audit-mode support, which is the only mode this project ever uses. tests/test_workflow_hardening.py's blanket "macOS is exempt" rule is narrowed to a named, ownership-scoped exemption for hardware-profile-smoke.yml only (owned by a different, concurrent task and off limits to edit here), with a new regression test enumerating this branch's three macOS jobs explicitly. 6. bootstrap-macos-brew no longer runs actions/setup-python (which could shadow python@3.13 with a GitHub-provided interpreter on PATH, making the check pass without ever exercising Homebrew's own install); it now checks a short, explicit list of brew --prefix candidates for python3.13 and fails with a clear message (not a false pass) if none resolve. 7. Path filters gained tests/test_adoption_*.py and the five scripts validate-macos runs. 8. The launchd smoke step's "qdrant binary absent" branch exited 0; qdrant is a pinned, always-selected component (not itself in required_commands, which is owned by adoption/manifest.json and not edited here), so this now exits 1 -- its absence after a successful bootstrap is a real failure to surface. 9. adoption/platforms/macos-arm64.md's two previously-cited green hosted runs are now marked as predating this fix round, naming the actual failing run (35812470345) and what it caught, so a reader does not treat stale hosted evidence as covering the current script. 10. brew_formulae's install order is now jq alone (needed by the --profile/pins validation that follows), then that validation, then the remaining five formulae -- previously all six installed before any validation ran. Evidence class: still no macOS host ran anything here. The npm/lockfile finding and the bash 3.2 abort are now independently reproduced (real npm, a real compiled bash 3.2.0, and a real /var-style symlink) rather than argued from memory; the launchd ownership/XML-escaping/harden-runner fixes remain offline local_integration and structural evidence pending the next hosted macos-15 run. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Rehash the evidence registry via host_receipts.register_file and rebuild the explorer Used scripts/host_receipts.py's register_file (an absolute Path.cwd() root, per the coordinator's fix-round instructions) for every file this fix round changed, instead of hand-patching manifests/evidence.json: the six already registered in the previous round's rehash (adoption-bootstrap.yml, bootstrap-macos.sh, pins-macos-arm64.json, macos-arm64.md, test_adoption_bootstrap_macos.py) plus scripts/adoption_status.py, and this round's four new files that were not previously registered (com.native-stack.qdrant.plist.template, launchd-agents.sh, test_adoption_launchd.py, test_workflow_hardening.py, render_launchd.py). register_file rewrites the whole manifest via json.dumps(indent=2), so this diff also reformats entries this round did not otherwise touch; the content change is exactly the updated/added hash and byte-count records. docs/ecosystem/index.html: regenerated with `scripts/build_ecosystem.py --write` from the manifests above and registered the same way; --check passes. python3 scripts/validate.py, scripts/validate_catalogs.py, scripts/validate_foundation.py --root . --json, scripts/landscape.py --root ., scripts/build_ecosystem.py --check and tools/sota-convergence/build_verdicts.py --check all pass after this commit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Round 2: defeat the npm shadow-copy, launchd orphan/unloaded/directory fixes, remaining lows Addresses the Opus round-2 re-review (fail): one High, one Medium, several Lows. Full evidence for each below. **High: npm's own unverified fetch shadowed the verified platform tarball.** install_npm installed the wrapper without deferring its own optional-dependency resolution, so npm auto-fetched the darwin-arm64 platform package itself, unverified, and nested it under the wrapper's own node_modules -- and Node's require.resolve checks that nested copy before a top-level sibling, so the previous round's `<alias>@file:<path>` top-level install was never what codex's bin/codex.js or claude-code's install.cjs actually resolved. Measured directly (npm 11.19.0) that neither --omit=optional, --no-optional nor NPM_CONFIG_OMIT=optional suppress the physical on-disk fetch for this install shape (a single global tarball install, no lockfile), and that pre-placing verified content at the nested path before installing the wrapper does not survive it either (npm overwrote it, "changed N packages"). install_platform_dependency now: installs the wrapper with --ignore-scripts first (deferring its lifecycle scripts); asks Node itself, via require.resolve scoped to the wrapper's own directory, exactly where it would resolve the dependency from; deletes that path and extracts the independently SHA-256-verified tarball there (falling back to a top-level alias only if Node found nothing to shadow); re-verifies resolution and the resolved version against the pin, fail closed on either mismatch; then install_npm runs `npm rebuild --global` for the wrapper, npm's own documented way to run the lifecycle scripts an --ignore-scripts install deferred, now against the verified copy. Proven end to end with real npm against two fixtures mimicking the two real shapes: a codex-like wrapper (no lifecycle scripts, resolves at every invocation) and a claude-code-like one (a postinstall that copies from wherever require.resolve finds the platform package into its own bin/ file, exactly like install.cjs) -- tests.PlatformDependencyInstallTests.test_codex_like_wrapper_resolves_the_verified_dependency_not_the_shadowed_one and .test_claude_code_like_wrapper_postinstall_copies_the_verified_dependency, each asserting the FINAL resolved/copied content is the verified fixture payload, never the differently-content "unverified" one an attacker-controlled registry fetch stands in for. **Medium: launchd-agents.sh install left an orphaned, unowned plist on a failed bootstrap.** cp ran before launchctl bootstrap; a failed bootstrap left the copied plist under ~/Library/LaunchAgents where RunAtLoad would reload it at the next login, and neither a retry (the ownership check) nor remove (never touches an unrecorded label) could ever reach it again. Now rolls the copy back (rm -f the destination) and exits 1 on a failed bootstrap, never recording it as enabled. **Lows:** - remove on an owned-but-unloaded label always failed calling bootout on something not loaded; now checks `launchctl print` first and, if not loaded, cleans up (deletes the plist, forgets the label) directly instead. - install's directory creation was hardcoded to this shell's own ECO_INSTALL_ROOT (state/logs, state/qdrant), which breaks a plist rendered with --host against a different host's ECO_ROOT value. Now reads each rendered plist's own StandardOutPath/StandardErrorPath/WorkingDirectory (via a new plist_value helper, plutil-or-plistlib-fallback like cmd_lint) and creates directories from those declared paths instead. - tests/test_workflow_hardening.py's MACOS_JOBS_OWNED_ELSEWHERE exempted the whole hardware-profile-smoke.yml file; keyed as "hardware-profile-smoke.yml:macos-profile" now, so that file's own ubuntu linux-profile job stays checked (new regression test confirms it does). - Path filters gained scripts/adoption_status.py in both push/pull_request lists. - render_launchd.py's render_plist now also catches ValueError from string.Template.substitute (a malformed placeholder, e.g. a bare trailing "$") and raises it as the same RenderError every other failure uses, instead of an uncaught traceback. - scripts/adoption_status.py's git_revision now also catches RuntimeError, which Path.resolve() raises on Python before 3.13 for a symlink loop (3.13+ instead returns the unresolved remainder); reproduced with a real two-node symlink loop on this host's Python 3.12. Also fixed in passing: tests/test_adoption_bootstrap.py's test_workflow_installs_the_manifest_supported_python_before_status did a whole-file text.index("scripts/adoption_status.py") search, which this round's new path-filter entry (a bare, unquoted-command list item earlier in the file than any job's real invocation) made match before the intended "actions/setup-python@" occurrence; narrowed the search to the literal "python3 scripts/adoption_status.py" invocation. Evidence class: the npm shadow-copy fix and the symlink-loop fix are real npm/Python behavior, measured and reproduced directly on this host (no macOS host ran any of this); the launchd fixes remain offline local_integration and structural evidence pending the next hosted macos-15 run. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Round 3: reinstall-while-loaded, arbitrary-path deletion, npmrc/postinstall gaps, atomic staging Addresses the Opus pass-with-findings review of round 2 (1c01f4f): the High (npm shadow-copy) is confirmed closed by that review (real tarballs downloaded, all four hashes confirmed, tarball contents and npm rebuild scope checked, prepare confirmed not to run). This commit fixes the remaining Medium/Low/Nit findings. 1. Medium, launchd-agents.sh install: re-running install on a label that is owned AND already loaded (e.g. after a re-render) had cp overwrite the live plist in place, then launchctl bootstrap fail (already bootstrapped), then the rollback delete the just-overwritten file out from under the still-running service. install now checks launchctl print first and boots the label out before reinstalling if it is loaded (refusing and leaving it running if that bootout itself fails); it only ever deletes the destination on a failed bootstrap when it did not already exist, otherwise it backs it up first and restores (and best-effort reloads) it. New shim test with a stateful launchctl (bootstrap sets a per-label loaded marker, print reads it, bootout clears it) drives install into exactly this already-loaded-reinstall-then-bootstrap-fails path and checks the original plist survives byte for byte. 2. Medium, bootstrap-macos.sh install_platform_dependency: `rm -rf` ran on whatever path require.resolve returned. Measured directly: Node's require.resolve, even given an explicit `paths` array, still searches its GLOBAL_FOLDERS fallback (NODE_PATH entries, $HOME/.node_modules, etc. -- Node's own module docs); setting NODE_PATH to a decoy directory made it resolve OUTSIDE the prefix entirely, and the previous code would have deleted that decoy -- a real, exploitable arbitrary-path deletion, not a theoretical one. The resolved path is now accepted only when it equals the exact expected nested path ($prefix/lib/node_modules/$package/node_modules/$dep_name/package.json); anything else falls back to the existing top-level alias instead of touching whatever require.resolve actually named. After installing the verified copy, the script now also asserts the resolved path equals $target_dir/package.json and that the installed package's own `name` field equals the pin's `resolved_package` (pinned since round 1, unread by any check until now), fail closed on either mismatch. New tests: a real decoy directory via NODE_PATH that survives completely untouched while the verified content still lands via the top-level fallback, and a resolved_package mismatch that fails closed even though the sha256 and version both verify. 3. Low: npm rebuild now passes --ignore-scripts=false explicitly, since a user-level .npmrc with ignore-scripts=true would otherwise make it return early without running anything. Added `postinstall_binary_check` (an optional pin field: platform_file / wrapper_file) and a cmp -s comparison after npm rebuild for claude-code specifically (its install.cjs copies package/claude, confirmed by downloading and `tar -tzf`ing the real 2.1.278 darwin-arm64 tarball, into bin/claude.exe): the check fails closed if what actually landed in bin/ does not byte-for-byte match the already-verified platform dependency, rather than trusting npm rebuild's own exit code alone. New tests cover both the matching case and a fabricated mismatch (a postinstall that writes fixed wrong content instead of resolving the platform dependency). 4. Low: a re-run of install_npm for the same id/version into an already-populated final_prefix would install directly into the live prefix, leaving a window where the wrapper's freshly updated files point at whatever platform dependency npm's own install just auto-fetched, unverified, before the fix-up runs. install_npm now stages into a fresh directory under stage_dir when a platform_dependency is involved and moves the finished, fully verified result into the final prefix only after install_platform_dependency succeeds (mv on the same filesystem is a single atomic rename), so the live prefix is always either the complete old install or the complete new one. New test runs install_npm twice for the same id/version and checks both succeed, the staging directory never survives either run, and the final content is still fully verified after the re-run. 5. Low, launchd-agents.sh remove: a launchctl print failure was treated as "not loaded" for ANY nonzero exit; now only 113 ("Could not find service") is treated that way. Any other print failure (permission, launchd itself unresponsive, ...) keeps ownership and reports, rather than guessing either way. The shared _launchctl_shim test helper was upgraded to track loaded state per label (bootstrap sets a marker, print reads it, bootout clears it) to match this more realistic behavior, and every affected test's expected launchctl-print/bootout counts were recomputed and verified against the actual, now-more-realistic call sequence. 6. Low: pins-macos-arm64.json's top-level `source` field and tests/test_adoption_bootstrap_macos.py's test_client_npm_pins_name_their_unpinned_darwin_arm64_dependency still described/named the round-1 `<alias>@file:<path>` design and the "unpinned" framing. Both updated; the test (renamed) now also asserts each platform_dependency carries a real sha256, not just a mention in install_note. 7. Nit: render_launchd.py's plistlib.dumps call moved inside the same try as _substitute, so a control character in a substituted value (which XML plist text content cannot represent) raises the same RenderError every other failure here does, instead of an uncaught ValueError. New test confirms this with a real control character. Evidence class: the shadow-copy defeat (closed by the coordinator's own independent Opus re-review, which downloaded and hashed the real tarballs itself), the NODE_PATH decoy, the resolved_package check and the atomic staging are all measured against real npm/Node on this host; the launchd ownership/reinstall fixes remain offline local_integration and structural evidence pending the next hosted macos-15 run (hosted CI at 522bbe3 is already green, 15/15, per the coordinator). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Regenerate the new-host grand list after #101/#102 and re-register round-3 evidence Reflects the macOS pin/socraticode changes in catalogs/landscape/new-host-grand-list.json and docs/new-host-grand-list.md (scripts/component_matrix.py --write produced no diff; scripts/new_host_grand_list.py --write did). Also re-registers the six files changed by round 3 (bootstrap-macos.sh, launchd-agents.sh, pins-macos-arm64.json, test_adoption_bootstrap_macos.py, test_adoption_launchd.py, render_launchd.py), which had been left registered at their pre-round-3 hashes after the origin/main merge. Verified via scripts/component_matrix.py --check, scripts/new_host_grand_list.py --check, scripts/evidence_manifest.py --check and scripts/validate.py, all passing. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * Round 3b: canonicalize paths across a symlinked ancestor; launchd install ordering/gating bootstrap-macos.sh (Opus M1 / hosted macos-15 run 35820422561; Codex round-3 non-atomic-swap Medium): - Add canonical_path(), a bash 3.2-safe, realpath(1)-free canonicalizer (python3 os.path.realpath, falling back to cd -P/pwd -P and a manual ancestor walk for a path that does not exist yet). A real Mac's /var is a symlink to /private/var, and Node's require.resolve realpath-resolves symlinks by default when it locates a module, so a prefix built under a symlinked ancestor (TMPDIR-based staging, or a test fixture root) never string-equaled the hand-built expected path. install_platform_dependency now canonicalizes its own prefix argument on entry and the resolved path Node reports; the post-install resolution assertion also realpaths both sides via fs.realpathSync. install_npm canonicalizes final_prefix and the staged/live prefix. Reproduced off-Mac with two new tests that build the fixture prefix/ecosystem_root under a symlinked tmp root (mirroring /var -> /private/var); both fail with "function canonical_path not found" (or, pre this fix, the exact "resolved to /priva..." mismatch) without the fix. - L1: replace the non-atomic `rm -rf "$final_prefix"; mv` prefix swap with rename-old-aside, move-staged-in, then delete-old -- a crash between the two never leaves final_prefix's own bin_dir symlinks pointing at nothing with the previous, fully verified install unrecoverably gone. - N2: rewrite install_platform_dependency's own design comment to state the exact-match containment condition and the resolved-package-name check, and the new canonicalization; update macos-arm64.md's platform_dependency paragraph (Codex), replacing its stale description of a lockfile/ `_integrity`-based verifier that no longer exists. launchd-agents.sh (Opus + Codex round-3): - L2: gate the pre-reinstall launchctl print/bootout probe on is_enabled_label, so install never unloads a label it does not own merely because it happens to be loaded. - L3: move directory creation and the owned-destination backup before that probe, so a failure in either never leaves an already-unloaded service down with nothing done to restore it. - L4: add wait_until_unloaded, a bounded poll of launchctl print (checked before any sleep, so the common case costs no wall-clock time) after a successful bootout, since bootout can return before real launchd teardown finishes; refuse rather than bootstrap over a service that never actually reported unloaded. - L5: install now treats only launchctl print exit 113 as confidently "not loaded", matching remove; any other failure refuses instead of silently assuming unloaded. - L6: add explicit remove coverage for print exiting 5 (keeps ownership and the plist; the existing code path was already correct, now asserted). - N1: move the temporary pre-overwrite backup out of ~/Library/LaunchAgents (which launchd auto-loads at login) into state_dir/backups, with a trap-based cleanup (pending_backup_plist) as the safety net for a premature exit. - Five new behavior tests plus three structural ones; three of the five (L2, L4, L5) fail against the pre-fix script, confirmed by stashing it. Full acceptance re-run on this commit: named macOS/launchd modules and the full suite (2688 tests) under bash 5, the same bash-3.2/launchd-gated subset (50 tests) under a real compiled GNU bash 3.2.0, --plan clean under both bash versions (shimmed Darwin/arm64), shellcheck and zizmor clean, git diff --check clean, scripts/validate.py and the catalog/foundation/component- matrix/grand-list/evidence-manifest checks all passing, and guarded gitleaks (dir scan and the git-history scan scoped to this branch's own commits since origin/main) finding no leaks. No macOS host ran any of this. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * Round 3c: rollback the failed prefix swap; stop the launchd trap truncating the live plist bootstrap-macos.sh (Codex round-3 verification of 5310fe5, Medium 1): - The rename-old-aside/move-staged-in swap had no rollback if the second move (staged prefix into final_prefix) itself failed: final_prefix was left absent with every bin_dir symlink into it broken, even though the previous, fully verified install still existed on disk as final_prefix.previous.$$. Failure injection (a shimmed mv failing on any "*-staged" source) confirmed this on 87327dc: exit 1, final_prefix absent, only the ".previous." copy remaining. Fixed by moving that previous copy straight back into final_prefix's place before returning, so a failed move loses only the newly staged prefix, never the working one bin_dir already pointed at. - Corrected the design comment's "atomic" claim (also flagged) to say what is actually guaranteed: this swap is recoverable, not atomic; a genuinely atomic swap would need a versioned directory plus a single symlink flip, a larger restructuring this fix does not make. - New failure-injection test proves the live prefix (and require.resolve's own resolution into it) survives a failed final move with the original content intact; confirmed failing against the pre-fix script by stashing it (reproduces the exact "absent, previous-only" outcome Codex reported). launchd-agents.sh (Codex round-3 verification of 5310fe5, Medium 2): - The round-3b N1 trap was itself a regression: `cp -- "$source_plist" "$dest_plist"` wrote directly into the live destination, so a copy that failed partway (disk full, an interrupted write) could truncate it and exit via `set -e` before any of cmd_install's own restore branches ran -- and the trap only DELETED the pending backup instead of restoring it, losing both the live plist's content and its one recovery copy (the pre-N1 script at least kept the backup in this scenario). Differential injection (a shimmed cp truncating any destination under .../LaunchAgents/*.plist) confirmed the regression against 87327dc's pre-3c script: destination truncated, backup deleted. - Fixed two ways: (1) install now copies to a same-directory "*.new.*" temp file and renames it into place, so a failed copy never reaches dest_plist at all; (2) the EXIT trap now restores pending_backup_dest from pending_backup_plist before removing the backup, rather than merely deleting it, covering whatever other failure might still reach it. - New failure-injection test (the same differential shim) proves the fixed script's reinstall succeeds cleanly -- the injection never fires at all, since the vulnerable direct-to-destination copy no longer exists -- and, confirmed by stashing the pre-fix script, that the same injection there reproduces the exact truncate-and-delete-the-backup regression. Full acceptance re-run: named macOS/launchd modules (115 tests) and the bash-3.2/launchd-gated subset under a real compiled GNU bash 3.2.0 all pass; bash -n and shellcheck clean on both scripts; scripts/validate.py and the evidence manifest check pass. No macOS host ran any of this. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * Re-register test_workflow_hardening.py after merging origin/main (#108) scripts/validate.py flagged a hash/byte-count mismatch for tests/test_workflow_hardening.py (part of #108's GitHub automation closure) immediately after merging origin/main; re-registered via host_receipts.register_file and re-normalized with evidence_manifest.py --write. All other post-merge checks (component_matrix, new_host_grand_list, evidence_manifest --check, validate_catalogs, validate_foundation, build_verdicts --check, the full 2796-test suite) pass. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * Round 3d: genuine recovery for both swaps (rollback, was_loaded/reload, atomic flip) Codex's round-3 verification of 5310fe5 returned pass-with-findings for everything except the two swaps' own recovery paths, which it reproduced by direct failure and signal injection. Fixed by design (step tables added as comments next to each swap; every row tested, including a failed rollback and SIGTERM mid-swap sent via kill -TERM $PPID from a shim, matching the coordinator's required method) rather than patched incrementally again: bootstrap-macos.sh -- replaced the rename-aside-then-move-in swap with the Homebrew Cellar/opt pattern: install into a versioned, never-reused directory (tools/<id>-<version>-<stamp>); final_prefix becomes a SYMLINK, flipped by a single python3 os.replace (genuinely atomic, unlike `ln -sfn`, which is unlink-then-symlink). bin_dir's own symlinks are unchanged -- they point into final_prefix/bin/* and transparently follow it through the extra indirection, so they never need re-creating on a later flip. The one non-atomic step this still has is a real, pre-3d final_prefix directory's one-time migration aside (rename() cannot replace a directory with a symlink in one call); the top-level cleanup() trap restores it if a failure or a signal lands in that window, reporting the exact manual `mv` if the restore itself also fails. Also fixed along the way: final_prefix must never be canonicalized (it silently resolved straight through the symlink to whatever it currently targets, defeating the whole design -- caught by the first new failure- injection test); pending_migration_prefix/_dest must be set BEFORE the migration mv, not after (a signal in the gap between the mv completing and the assignment left the trap unable to find what to restore -- caught by the SIGTERM test); the test harness (_run_install_npm) needed the same top-level cleanup()/EXIT trap the real script installs, which it had never included, so failure/signal recovery was silently never exercised at all until this was added. launchd-agents.sh -- the EXIT trap now restores only by same-directory rename (never `cp`, whose own failure was previously ignored before deleting the backup unconditionally); deletes the backup only after that rename succeeds, otherwise keeps it and reports the exact manual command; records was_loaded (via pending_backup_reload_needed, set right after a confirmed bootout) so restoring the backup from ANY exit path -- not just the explicit bootstrap-failure branch -- also reloads the service, not just its file. cmd_install's own explicit restore branch was removed entirely in favor of the trap (round 3c/3b had two independent implementations of the same recovery that had independently drifted different bugs; one is now authoritative). The backup itself moved back into dest_plist's own directory (reverting round 3b's N1 relocation under state_dir): `mv` across a filesystem boundary silently falls back to a non-atomic copy-then-unlink, so the restore-by-rename this fix depends on is only genuinely atomic when the backup is guaranteed to be on the same filesystem as dest_plist, which only its own directory can guarantee. New failure/signal-injection tests (7 total: 3 for the atomic flip, 4 for launchd's trap), each confirmed to fail against 96f1abe by stashing it. Full acceptance re-run on this commit: full suite (2801 tests) passes except the one evidence-registration lag scripts/validate.py itself always flags immediately after an edit (resolved by the registration commit that follows); named macOS/launchd modules (120 tests) and the bash-3.2/ launchd-gated subset (54 tests) under a real compiled GNU bash 3.2.0 all pass; --plan clean under both bash versions; shellcheck and zizmor clean; git diff --check clean. No macOS host ran any of this. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * Round 3e: idempotent convergence for launchd, ownership-checked pruning for bootstrap Codex's round-3e review of c18aea3 found two new Highs and four Mediums -- each earlier round's hand-written rollback had closed one edge case and opened another (an unowned rm -rf, a partial hard-copy backup, a reload skipped at a step boundary, cleanup left armed after success, a signal racing a foreground mv, an unguarded rm aborting recovery under set -e). Per the coordinator's explicit direction, this round changes the approach instead of patching further: simplify toward idempotent convergence, the way `brew services` works, with step tables and signal-injection tests proving every row. Both scripts now explicitly trap INT, TERM and HUP (`exit 130/143/129` only; recovery logic stays in the EXIT handler alone) -- bash only reliably defers a caught signal to a completed step when the signal is explicitly trapped, not left at its default disposition (confirmed directly: an untrapped SIGINT sent to a bash script with only an EXIT trap registered does not interrupt it at all in this environment, while an untrapped SIGTERM does -- the asymmetry Codex's finding 5 was pointing at). bootstrap-macos.sh (High #1, unowned rm -rf): new prune_old_version() deletes a superseded versioned directory only when its canonical form is a direct child of the canonical tools/ directory with a name matching <id>-<version>-<stamp>; an external target, a relative "../" target, a non-matching name, or a symlink loop (canonical_path's python3 os.path.realpath handles cycles without hanging) is logged and left in place, never deleted. Pruning is now genuinely best-effort: a rejected or failed deletion warns and never fails the install. Four new tests, one per adversarial target shape, confirmed to fail against c18aea3 (the function does not exist there). launchd-agents.sh -- full redesign, net roughly flat in size (+22 lines) despite ~390 changed, since one pure function replaces four tracking markers (pending_backup_plist/_dest/_reload_needed, pending_dest_tmp) and their two independent, historically drifting recovery implementations: - backup_plist is a hard link (`ln`, High #2), not `cp`: it either exists completely or not at all, so a partial copy can never overwrite the intact original. Same directory as dest_plist (round 3d), never this script's own state_dir. - Staging (cp to dest_plist.new, then mv over dest_plist) is unchanged in shape but the name lost its PID suffix, matching the backup's own naming, so a stale artifact from a crashed run is visible to and reconciled by the very next run, not just the same process. - reconcile_install replaces every marker: it reads dest_plist's identity against backup_plist (-ef, since the backup is a hard link) and launchd's own current loaded state, and converges to whichever ONE action that state calls for (drop the backup, reload, restore-and- reload, or -- if nothing converges -- keep the backup and name the exact re-run command). It is the ONLY recovery logic: called explicitly after a normal attempt, and identically as the EXIT trap for anything that interrupts one. `set +e` inside it, explicit checks throughout, per the coordinator's instruction that the handler must never itself become fatal under the script's own `set -e`. - record_enabled_label now runs as soon as install actually commits to a label (not only on success), so a label that ends up needing attention is never re-refused as unowned on retry; a was_already_enabled snapshot, taken before that point, keeps the L2 bootout gate correctly checking ownership as it stood BEFORE this run, not after. New tests: a shared _signal_shim helper plus a parameterized step x signal matrix (ln and mv, TERM and INT) proving retry-converges for launchd; a dedicated SIGTERM/SIGINT-parameterized migration test and three more for bootstrap-macos.sh's prune_old_version. Every new/changed test confirmed against c18aea3 by stashing the fix; several rewritten round-3b/c/d tests (now asserting the new design's own honest behavior -- e.g. a fresh install that never loads is left on disk with a manual-retry message, never silently deleted) checked to still exercise the coordinator's original findings under the new mechanism. Full acceptance re-run on this commit: full suite (2821 tests) passes except the one evidence-registration lag scripts/validate.py itself always flags immediately after an edit (resolved by the registration commit that follows); named macOS/launchd modules (126 tests) and the bash-3.2/ launchd-gated subset (56 tests) under a real compiled GNU bash 3.2.0 all pass; --plan clean under both bash versions; shellcheck and zizmor clean; git diff --check clean; guarded gitleaks clean. origin/main had not moved since it was already merged into c18aea3's own parent, so no new merge was needed this round. No macOS host ran any of this. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * Round 3f: abort preflight on unresolved reconcile, path-verify launchctl state, poll through delayed teardown Codex verified d51f062 FAIL with a High and two Medium regressions in launchd-agents.sh's reconcile_install/cmd_install, plus an open delayed- teardown boundary; bootstrap-macos.sh itself was confirmed sound and is untouched here (its test matrix gained the same row/signal coverage). - High: reconcile_install's own preflight failure is no longer ignored (`|| exit 1`, not `|| true`): a retry that still cannot converge now aborts before ever reaching the backup-overwriting `ln`, so the last known-good backup can no longer be silently destroyed and re-linked to unusable content. - Medium (tri-state print): launchctl_label_state() reports loaded_here, loaded_elsewhere, not_found or unknown instead of a binary loaded check; unknown now takes no destructive action anywhere and keeps the backup. - Medium (unrelated loaded label): ownership is recorded only inside reconcile_install, once a path match confirms dest_plist itself is what launchd has loaded -- never earlier -- and both reconcile_install's fresh-install path and cmd_install's own bootout gate refuse outright on loaded_elsewhere, on the first invocation and every retry. - Delayed teardown: reconcile_install re-polls (reusing wait_until_unloaded) when this run's own bootout is still pending before trusting a "loaded" read, instead of declaring convergence with zero reload attempts. Added four named regression tests reproducing Codex's three exact scenarios plus the delayed-teardown case, each verified via differential (git stash) to fail on d51f062 and pass here. Expanded the checked-in matrix to cover every row of both step tables under failure, SIGTERM and SIGINT (was two launchd rows and one bootstrap row), reusing and extending the existing signal-shim helpers rather than one bespoke script per case. Updated the in-script step table for the ownership-timing and preflight- abort changes, and recorded the delayed-teardown re-poll as an untested boundary for real-Mac acceptance in the platform notes: modelled against a mock launchctl and real bash signal injection, never observed on native launchd. Verified with the full adoption test suite (61 launchd + 78 bootstrap-macos tests) under both bash 5.2 and a real compiled bash 3.2.0, shellcheck, zizmor, a manual bash-3.2 offline render/lint/install/status/ remove cycle, a manual bash-3.2 real-SIGTERM injection during install, scripts/validate.py and friends, and guarded gitleaks (dir + git log range) -- all clean. origin/main has not moved past this branch's base; no merge needed. Evidence class unchanged: Linux-side local_integration/structural plus real npm/Node/bash-3.2/signal behavior measured directly on this host; no macOS host has run any of this. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * Scope the pre-reinstall load-check test to cmd_install's own body test_install_gates_its_pre_reinstall_load_check_on_ownership matched cmd_remove's textually similar is_enabled_label + launchctl print gate instead of cmd_install's round-3f one, so it passed for the wrong reason. Added _shell_functions (the same top-level-function extraction helper tests/test_adoption_bootstrap_macos.py already uses) to scope the assertion to cmd_install's own body, and rewrote it to check the round-3f gate directly: launchctl_label_state is called before any bootout, and the loaded_elsewhere and unknown case arms both refuse (exit 1, no bootout) rather than proceed. Verified the new test actually exercises cmd_install and not cmd_remove by temporarily replacing cmd_install's gate with the old pre-round-3f binary launchctl-print check (no path verification, no loaded_elsewhere/unknown refusal) and confirming the test fails against that mutation, then restoring the gate byte-for-byte (diff against HEAD is empty for launchd-agents.sh; only the test file and evidence manifest changed here). Re-ran the full test_adoption_launchd suite (61/61) under bash 5.2 and a real compiled bash 3.2.0, plus bash -n on both interpreters, shellcheck, scripts/validate.py, evidence_manifest.py --check (re-registered via host_receipts.register_file) and guarded gitleaks dir -- all clean. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * Round 3g: replace launchd's transactional backup/reconcile with brew-services semantics Codex's review of 0e1e0d4 found four new Medium regressions in the round 3b-3f backup/reconcile design (remove booting out an unrelated service on stale state; a print error on fresh install deleting an already-loaded plist; a signal during bootout or restore leaving the service stopped with no retry built in; an ownership-write failure falsely reporting success) -- the fifth straight round to find a new recovery defect in that machinery. Rather than patch a sixth, launchd-agents.sh now converges on the maintained upstream pattern: Homebrew/brew Library/Homebrew/services/ cli.rb @ 8e3a5dc0a7 (2026-09-07), which has no backup, no rollback and no ownership file, and recovers by re-running the command. - Removed record_enabled_label, is_enabled_label, forget_enabled_label, the enabled-labels state file, reconcile_install, was_already_enabled and the hard-linked backup entirely. - Stateless ownership: "ours" means the label is one of this script's own AND, if it is loaded at all, launchctl print's own "path =" line names this script's destination plist -- never a stored record. install and remove share the exact same launchctl_label_state check (tri-state plus elsewhere, unchanged from round 3f) and both refuse outright, touching nothing, on loaded_elsewhere or unknown. - install: stage a temp file beside the destination and lint it; if loaded_here, bootout and wait until confirmed unloaded (refusing otherwise); rename into place; launchctl enable + bootstrap; on bootstrap failure, leave the new plist in place (no rollback -- there is nothing to roll back to) and print the exact recovery commands. - remove: on loaded_here, bootout, wait, then delete; on not_found, delete directly if a file is present; refuse on loaded_elsewhere or unknown. - The EXIT trap now only removes the current run's own still-inert temp file; INT/TERM/HUP stay explicitly trapped so bash always defers to a completed step. Replaced every round 3b-3f backup/reconcile test with a convergence matrix (cp/bootout/mv/bootstrap x {failure, TERM, INT}, then a fault-free install retry converging to NEW loaded, then a fault-free remove leaving nothing loaded and no file) plus named tests for loaded_elsewhere refused by both install and remove, unknown touching nothing on both, and Codex's exact remove reproduction (a previously installed label's plist on disk, its label now loaded from a different, unrelated path). All new/changed tests verified via differential (git stash on the script alone) to fail on 0e1e0d4 (19 of 55 tests fail there, the rest being unaffected render/ template/structural checks); all 55 pass here. adoption/platforms/macos-arm64.md carries the dated decision (chosen: this design, with the Homebrew citation; rejected: the transactional design, because rounds 3b-3f each found a new recovery defect in it; would overturn: a real-Mac observation where a re-run does not converge, or a requirement to preserve a locally edited plist) and keeps the untested delayed-teardown-polling boundary note, updated for the new call sites. launchd-agents.sh shrank from 758 to 593 lines (38310 to 28554 bytes) against 0e1e0d4, removing the transactional machinery outright rather than adding another layer of recovery logic on top of it. Merged origin/main (bb5228d, unrelated: release-tag pin bump) -- no conflicts with this file set. Verified: full test_adoption_launchd (55/55) and test_adoption_bootstrap_macos (78/78, unaffected) under bash 5.2.21 and a real compiled bash 3.2.0; a manual bash-3.2 offline render/install/ status/remove cycle and a manual bash-3.2 real-SIGTERM injection during the rename step; shellcheck, bash -n on both interpreters, zizmor, git diff --check, scripts/validate.py and friends (evidence re-registered via host_receipts.register_file), and guarded gitleaks (dir) -- all clean. Evidence class unchanged: Linux-side local_integration/structural plus real bash-3.2/signal behavior measured directly on this host; no macOS host has run any of this. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * Round 3h: pin and download the llama-embed model, add the embedding acceptance check A 2026-09-23 readiness audit found two defects in PR #94: 1. The llama-embed launchd agent's template ran llama-server with no model argument at all -- KeepAlive would restart it in a loop forever on a real Mac, never actually serving embeddings. Fixed: - com.native-stack.llama-embed.plist.template now passes -m ${EMBED_MODEL_PATH}. - adoption/pins-macos-arm64.json gained a separate top-level "models" section (never "tools", so it is never swept into adoption/manifest.json's protected profile-component coverage) for ggml-org/embeddinggemma-300M-GGUF @ 0f741b5a6585bd53aeb15cd1372c56f2a0f65e12, embeddinggemma-300M-Q8_0.gguf, sha256-pinned, with its own huggingface_lfs_oid_plus_local_rehash checksum source. - adoption/bootstrap-macos.sh gained install_embed_model, a dedicated download-and-verify step (reusing fetch(), the same helper every other pin already uses) run unconditionally after the tools[] loop, fail-closed on a null or mismatched sha256, into $ECO_INSTALL_ROOT/state/models/. - launchd-agents.sh's cmd_render and adoption/hosts/macos-example.json both carry EMBED_MODEL_PATH now, alongside the existing HOME/ ECO_ROOT/AI_MEMORY_URL host values. - New tests: the rendered plist's -m argument matches EMBED_MODEL_PATH (test_adoption_launchd.py); install_embed_model, extracted verbatim and run against a real (curl-shimmed) fetch(), installs on a matching digest and fails closed on a mismatch or a null sha256 (test_adoption_bootstrap_macos.py); the models[] pin's own shape, URL-matches-revision and separation from tools[]. 2. No Linux reference vector existed to compare a macOS response against. Fixed: - evidence/artifacts/macos-embed-reference-20260923/ carries the captured reference (llama.cpp b11057 ubuntu-x64 CPU build, the identical pinned GGUF, 768 dimensions, L2-normalized, repeat cosine 1.0, full request/response provenance), registered via host_receipts.register_file. - tools/adoption/embed_acceptance.py (stdlib only) sends the reference's exact request body to a running llama-server, checks dimension and cosine >= threshold against the reference embedding, and prints a JSON result; understands both the OpenAI-compatible and llama.cpp-native response shapes, and reports a network or malformed-response failure as JSON, never a bare traceback. Six new tests (a local stdlib HTTP server stands in for llama-server) cover an exact match, both response shapes, a dimension mismatch, a below-threshold cosine, and a connection failure. - .github/workflows/adoption-bootstrap.yml's bootstrap-macos job now caches the pinned model (actions/cache, keyed on its sha256, not a version tag), installs and health-checks the llama-embed launchd agent, runs embed_acceptance.py against it, uploads the JSON result and the startup log, and best-effort-greps the log for which backend it reports -- labelled explicitly as hosted-runner evidence only (a GitHub-hosted macOS runner is CPU-only, never a claim of Metal actually engaging; a real Mac run is what would show that). - adoption/platforms/macos-arm64.md's "What a hosted run proves" and "Acceptance test for the 24 GB default" sections both updated: the reference vector now exists, and the one command to run it is documented, still bounded to "unrun on any Mac, hosted or otherwise" for the acceptance itself. Verified: full test_adoption_bootstrap_macos (93/93) and test_adoption_launchd (56/56) under bash 5.2.21 and a real compiled bash 3.2.0; bash -n and shellcheck clean on both scripts; zizmor clean on the workflow; git diff --check; scripts/validate.py and friends (evidence re-registered via host_receipts.register_file); guarded gitleaks dir -- all clean. Evidence class unchanged and explicit throughout: Linux-side local_integration/structural, a real captured Linux reference vector, and (once this round's hosted CI step runs) CPU-only hosted-runner evidence; no macOS workstation has run any of this. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * Round 3i: canonical-path comparison, EINPROGRESS handling, cache-before-verify, embed_acceptance transport errors, macOS recording-tooling CI Codex's review of e9334e9 returned FAIL with three Medium findings and one Low, all with narrow, well-isolated fixes: 1. launchd-agents.sh's loaded-state classifier used `-ef` (device+inode identity), so a DISTINCT hard-linked filename sharing dest_plist's own inode incorrectly compared loaded_here, letting remove bootout and delete a plist for a label actually loaded from an unrelated, hard-linked path. Replaced with canonical_plist_path: resolves symlinks in a path's parent directory only (cd -P, falling back to python3's os.path.realpath), then reattaches the final component unresolved -- neither mechanism ever collapses a hard link's own distinct pathname the way inode identity did, since a hard link has no stored "canonical name" to resolve to. 2. A `launchctl bootout` returning EINPROGRESS (teardown already under way, from this run's own retry of an earlier interrupted attempt, or another process) was treated as an unconditional failure, abandoning an in-progress teardown with no wait and no reload ever attempted. The pinned Homebrew cli.rb @ 8e3a5dc0a7 (services/cli.rb, ~line 311) retries on exactly Errno::EINPROGRESS (Darwin 36) for the identical reason. New shared bootout_and_wait absorbs it the same way, then polls (reusing wait_until_unloaded) rather than refusing outright. Self-discovered along the way: the first implementation captured its result via `result="$(bootout_and_wait ...)"`, which forks a subshell -- empirically verified that a signal delivered while a foreign command runs INSIDE that subshell is silently swallowed there and never reaches this script's own INT/TERM/HUP traps at all. Fixed by using a plain function call plus a global result variable instead, keeping bootout_and_wait in the same process as its caller. 3. adoption-bootstrap.yml's embedding-model cache restore ran AFTER bootstrap's own checksum verification and used a literal hardcoded sha256 as its key. A restore running that late could silently overwrite the just-verified model with stale cached bytes, with nothing left to re-check them, and a changed pin would keep matching the same stale cache entry. Fixed: a new step reads the pin's sha256 into a step output, the cache restore (now keyed on that output) moves to before the bootstrap step, and bootstrap-macos.sh's own fetch() (unchanged, already correct) transparently re-downloads and re-verifies any mismatched destination on the exact same path a genuine cache miss takes. 4. embed_acceptance.py caught urllib.error.URLError around the request, but response.read() can raise a bare TimeoutError (a stall mid-body, past urlopen's own connect-phase timeout) or http.client. IncompleteRead (the server closing before its declared Content-Length was satisfied) -- neither caught, leaving stdout empty and the process exiting via an uncaught traceback instead of the documented JSON failure contract. Both now caught explicitly, plus a catch-all OSError. Also, this round (peer-update-audit gap adoption_macos, medium): the catalog's own recording and verdict scripts (host_receipts.py, component_matrix.py, new_host_grand_list.py, build_verdicts.py, validate_convergence.py, release_due.py) had never run on macOS CI or against macOS's own system Python -- every existing gating job for them runs on ubuntu-24.04 only. adoption-bootstrap.yml's validate-macos job now gates on all six, then runs a recording smoke: installs the profile, records a real host_receipts.py record receipt (component codex, --from-stack-commands, --evidence-class native_proven -- hosted-runner evidence only, never a workstation acceptance receipt) inside a throwaway copy of the checkout ($RUNNER_TEMP/rec, never the real one, so nothing is committed), then re-validates and re-checks the matrix/grand-list in that same copy. Runs once against the manifest-pinned Python line and once against the runner's own system /usr/bin/python3 if it meets the newly declared Python 3.9 floor (adoption/bootstrap.md, adoption/platforms/ macos-arm64.md's new "Recording and verdict scripts" section), skipping with a message otherwise. Added named tests for all five changes (a hard-linked alias never compares loaded_here; a retry converges while an earlier run's teardown is still asynchronously in progress, fixing the bootout-signal test shim Codex flagged for clearing the marker before signalling and so hiding the async case; CI step order and the derived cache key; two embed_acceptance transport-failure modes over a raw-socket server; structural checks on the new validate-macos steps), each verified via differential (git stash) to fail on e9334e9 or the pre-fix file. Full suites: test_adoption_launchd 59/59, test_adoption_bootstrap_macos 107/107. Verified: both scripts under bash 5.2.21 and a real compiled bash 3.2.0 (syntax and shellcheck), PLUS two direct, real-SIGTERM/real-shim reproductions under bash 3.2 outside the Python harness (the hard-link refusal, and the EINPROGRESS retry actually converging); zizmor and actionlint both clean on the workflow; git diff --check; scripts/ validate.py and friends including host_receipts.py validate and build_verdicts.py --check (evidence re-registered via host_receipts. register_file); guarded gitleaks dir -- all clean. origin/main has not moved past this branch's base; no merge needed. Evidence class unchanged: Linux-side local_integration/structural plus real bash-3.2/signal behavior measured directly on this host; the new CI steps are themselves unrun until a real macos-15 job executes them. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * Fetch full history in validate-macos so host receipt catalog revisions resolve Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Resolve symlinked plist leaves in canonical_plist_path; make the incomplete-body test drain the request canonical_plist_path now follows a symlinked leaf (bounded at 40 hops) so a plist launchctl reports by its resolved target compares loaded_here, while a hard link stays a distinct path (Codex round-3i finding 1). The raw-socket test server drains the whole request and half-closes, so the client sees IncompleteRead rather than an occasional ECONNRESET (30/30 runs pass). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Skip the PyYAML structural workflow tests where PyYAML is absent (repository convention) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Give the macOS recording smoke a per-run --identity (required since #117) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Record the 2026-09-23 hosted macOS smoke receipt (launchd, embedding acceptance, recording smoke) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Round 3j: fix 8 unresolved Codex-connector review threads blocking the #94 merge 1. launchd-agents.sh:77 -- llama-embed was left out of the default install/ status/remove label set, even though bootstrap now downloads and sha256-verifies its model unconditionally (round 3h). Added it to default_labels; --label was the only way to reach it before this. 2. launchd-agents.sh:526 -- install created the plist's own declared StandardOutPath/StandardErrorPath/WorkingDirectory BEFORE the ownership check, so a refused install (loaded_elsewhere or unknown, which must touch nothing) still left new, empty directories behind. Moved directory creation to after the ownership check's case statement. 3. macos-arm64.md:217 -- "What a hosted run proves" was stale (cited a 2026-09-22 run, claimed launchd and the embedding acceptance were both still unrun). Rewritten for run 35875188590 at head 75a6e0d: launchd genuinely bootstrapped and booted out both qdrant and llama-embed; qdrant health reported 1.19.1; the embedding acceptance passed with cosine 0.99944 and the startup log reports Metal device MTL0 engaging for the layers it covers; the recording smoke passed on both actions/setup-python 3.13.15 and the runner's own system /usr/bin/python3 3.9.6. Kept the hosted-runner-only scope throughout -- no platform_status change, no workstation acceptance claim. 4. qdrant.plist.template:28 (the most important one) -- nothing on a clean host ever created $ECO_ROOT/config/qdrant.yaml; only the hosted CI step's own disposable file did. Added provision_qdrant_config to bootstrap-macos.sh: mirrors examples/qdrant.yaml.example exactly (storage under $ECO_ROOT/state/qdrant, loopback, port 16333 -- the same port adoption/hosts/macos-example.json's QDRANT_URL and recipes/README.md both name), plan-mode aware, and never overwrites an existing file. The hosted launchd step no longer writes its own disposable config or uses port 6333; it uses the provisioned one at port 16333, exactly what a real install would use. 5. ai-memory.plist.template:30 -- added --workspace local --project native-agent-stack, matching the selected native recipe's own serve invocation byte for byte (recipes/README.md:244; the same scope docs/foundation-stack.md:51-52 and every other documented ai-memory invocation in this project already use). 6. adoption-bootstrap.yml:13 -- the paths triggers covered adoption/** and tools/adoption/** but missed most of what validate-macos's recording-tooling gate and recording smoke actually consume. Added the embed reference artifact, host_receipts.py, component_matrix.py, new_host_grand_list.py, build_verdicts.py, platform_status.py, validate_convergence.py, release_due.py and manifests/evidence.json to both push and pull_request path lists. 7. launchd-agents.sh:517 -- lint only checked plist syntax, never that a rendered plist's own Label key matches its filename (also the label every other subcommand addresses it by, and what launchd itself expects to agree). Checked once syntax lint has already passed, via pure bash parameter expansion (no new basename/dirname dependency). 8. launchd-agents.sh:620 -- status always exited 0, even when every requested label's launchctl print failed (absent, 113, or launchd itself unavailable, any other nonzero exit) -- a caller checking only the exit code never saw either case. Now returns non-zero whenever any requested label's print fails…
9 of 22 tasks
seathatflowsinourveins
deleted the
codex/evidence-convergence-practice
branch
September 25, 2026 18:47
This was referenced Sep 27, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This adds a reusable research-to-acceptance practice and two executed retrieval comparisons to the native stack. Public repository, star and curated-list discovery now connects to pinned source review, frozen experiments, observed results and explicit adoption gates across six capability areas.
The source wave refreshes one public owned repository and all 342 public stars, inspects selected links in three pinned awesome lists, and reviews seven candidates. Registered decisions extend the typed union to 504 repository identities. Primary source and license hashes are retained separately from execution evidence.
The exact-release Agent Retrieval Bench replay covers all 101
trace2codecases with zero skips. Its lexical baseline has higher mean Recall@20 than BM25 on that subset; a separate frozen 24-question public-document fixture favors BM25 at Recall@1 on its 20 answerable cases. The four no-answer cases retain rankings without inventing an abstention score. Both packets include exact pins, hashes, per-case results, replay instructions and limits; they establish retrieval behavior rather than agent repair, model savings or host-wide adoption.Validation:
The change targets current
main, preserving its Mac worker fix and newer security-identity research. All original convergence source and result artifacts remain unchanged.