Repository navigation
Stage 1 client-layer evidence: claude-code and codex receipts on mac-coordinator-64gb-20260925 (#382) - #391
Conversation
…b-20260925 (#382) Records native_proven use-stage receipts for the installed claude-code and codex clients on the Mac's Stage 1 client layer (host request #382): a version-resolving command, a headless claude -p turn with a random-fixture positive control plus claude mcp list, and a codex exec turn (JSON event stream, command_execution positive control) plus codex mcp list --json and codex features list cross-checked against this host's own config.toml. Both receipts are recorded with --allow-unbound-version (installed builds are newer than the current catalog pins) and superseded once each to back the version claim with a retained command instead of only --component-version. Refreshes the derived component-evidence-matrix accordingly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ent-layer-20260927
Merges origin/main (no conflicts) and re-registers the derived matrix/grand- list files. Adds evidence/artifacts/mac-stage1-client-layer-20260927/README.md: the coordinator's Stage 1 client-layer install on mac-coordinator-64gb-20260925 (backup summary and rollback, launchctl before/after plus a fresh live re-check, the headless read-back names/counts, plugin revision check, skills status, token-efficiency/client-wiring coverage, the Codex profile, and the macOS template findings), each line marked as coordinator-reported or independently re-checked in this session. Notes two additional findings this session found while recording the codex receipt: the installed codex-cli routes through the ChatGPT desktop app's bundled build, and codex exec needs stdin redirected from /dev/null or it blocks indefinitely. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…TH inference Reads ~/.claude/settings.json's own env.PATH value and re-runs command -v rtk / rtk --version with exactly that PATH (the environment Claude Code gives its hooks) instead of inferring resolution order from the PATH list. Confirms it resolves to the ecosystem install and reports rtk 0.50.0. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ent-layer-20260927
…cord receipts with verified checks, disclose Codex/cross-host surface Applies the round of review findings against PR #391: - Record -3 generations of both host receipts, superseding -2: fix the codex version-check command's exit-code bug (a later echo, not `codex --version`, decided it), capture the model's actual command_execution text instead of assuming `wc -l`, disclose live-only MCP servers (cua_repl, codex_app) without failing the receipt on them, drop the unchecked approval_policy/ sandbox_mode assertion, add a Read-tool-use check and a real claude mcp list exit-code/status gate to the claude-code receipt, and add an in-transcript negative control to both. - Relabel three PR-body evidence-class rows from native_proven to source_review where no retained command backs the claim (launchctl PID-continuity comparison to the coordinator's private captures, the Codex profile's approval_policy/sandbox_mode, and the rtk-hook/Qdrant-URL "independently reconfirmed" line), and remove a corroboration claim that shared its source with what it claimed to corroborate. - README: disclose the Codex live MCP surface (10 servers vs. 8 declared in config.toml, the two extras plugin-provided) and the applied remoteControlAtStartup/crossSessionInbound settings, their user scope and interaction with this profile's bypassPermissions default; correct the backup/rollback section to state what is and is not confirmed and to actually undo Stage 1's additions instead of merging over them; record the plugin-revision check's installed SHAs and GitHub compare API output; state that a compliant independent review of the -3 receipts cannot return agree (both are --allow-unbound-version). - Open #394 for the two pre-existing ecosystem agents' tools-allowlist gaps (host state, not in this repository's diff). - Fix a lanes.md wording nit: manifests/evidence.json is a shared hot file touched for registration only, not foundation-owned. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…ent-layer-20260927
…re source_review Both claims were checked directly in this session, not through either host receipt's own retained `commands` (neither's command array prints an rtk version or a server address), so they were still labeled `native_proven` in the README even after the prior review-fix pass (b521b79), which relabeled three other rows the same way but left this exact fragment ("rtk-hook-wiring and the Qdrant URL were independently reconfirmed native_proven") unchanged. Confirmed against both -3 receipts' full command arrays: neither runs an rtk command, and the codex mcp-list command's own filter only ever prints name/enabled/transport, never a server URL, by design (to avoid capturing MCP server config in the receipt). README: both bullets now state the claim is source_review, matching the wording already used for the Codex profile and launchctl-continuity claims in the same document. The PR body's evidence-class table gets the same relabel via gh pr edit (not a commit), alongside refreshed base-commit/ verified-file-list/head references for the origin/main merge below. Also completes the hot-file merge this commit's parent started: took main's manifests/evidence.json (2 new commits landed upstream: 805991b, 6c073fc) rather than trust the clean auto-merge, and re-registered only this branch's 7 files (6 receipts + README) through host_receipts.py's register_file. Re-verified: python3 scripts/host_receipts.py validate, scripts/component_matrix.py --write, scripts/new_host_grand_list.py --write, scripts/validate.py (FAILS=0), git diff --check, gitleaks over both touched evidence directories (0 leaks), and the three targeted unittest modules (tests.test_host_receipts, tests.test_component_matrix, tests.test_new_host_grand_list -- 261 tests, OK). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
Review findings: full dispositionBoth saved Opus reviews ( review-evidence lens (5 should_fix, 8 nit)
review-security lens (3 should_fix, 5 nit, 1 no-findings)
What's outstanding (not fixable from this PR alone)
Hot-file protocol
CI (at time of this comment)18 checks passing, 3 skipped (expected), 3 in progress: |
…ent-layer-20260927
…ent-layer-20260927
…ry writes, record codex -4 with a fixture-gated positive control Rewrites the README rollback so it can only add or overwrite files the private backup holds (rsync/cp, no --delete, no rm) and lists every Stage 1 addition it does not remove with its own native removal command. Discloses that every receipt's model turn fires the live ai-memory hooks and writes to the running production store, and that the declared context-mode MCP server and Codex's plugin-provided tools both run outside --sandbox read-only. Records mac-coordinator-64gb-20260925--codex--use--20260927-4.json (supersedes -3): its positive control now gates on and prints a command_execution whose command text actually references the fixture, rather than the first successful command_execution of any kind; also fixes mcp list/features list to use set -o pipefail with stderr discarded before the parse. Relabels the MCP-surface provenance sub-claim source_review instead of native_proven in the README and PR body. Fixes wording that treated the coordinator session and this evidence-PR session as independent observers -- every receipt in this PR carries the identical recorded_by.identity_sha256, so they are the same session -- and reroutes the same-host independent review to a fresh session accordingly. Plus the cross-reference, excerpt-description, pin-wording and merge-count nits from the re-review. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
Review-fix pass, part 3: addressed both lenses' re-review (full disposition in the PR body's new "Review-fix pass, part 3" section).
Checks (all from the worktree at the new head):
New head: |
…ent-layer-20260927
… ordering and receipt citations Fixes nine defects from an independent Opus re-review of the Stage 1 client-layer README: - Replaces the unscoped `<skills-bin> remove --all` with one `skills remove <name> -g -y` per pinned skill (all 28), citing the decision record's -g scoping rule and lifecycle.md's "never uninstall all user tools to roll back one package". - Replaces the directory-wide workflows delete with a name-and-sha256 verified per-file removal of the 13 files Stage 1 actually placed in ~/.claude/workflows (12 byte-identical to the repo source, one a per-host template render). - Reorders the rollback so every removal runs before the restore -- a restore run first would overwrite the live state the removal table was built from -- and rescopes the "never runs rm" claim to the restore step alone, since removal now legitimately runs rm. - Joins the one-line restore's three commands with `;` instead of `&&` so a missing claude.json backup can't silently skip the codex rsync. - Documents that only two of the four hook files install_claude_profile.py's HOOKS map can place actually exist on this host; the other two were added to that map after Stage 1's own install ran here. - Replaces the blind rm of app-server-daemon/settings.json with a conditional that restores from the private backup if it holds a copy, otherwise removes only the one key this session is known to have written, since the file's prior existence is unrecorded. - Corrects the claim that both current receipts print a sanitized command line: only codex -4 does; claude-code -3 prints tool names only. - Fixes a section citation (contributing-evidence.md section 3 step 8, not section 8). - Adds a sentence labeling the -4 receipt's plugin-to-server attribution (line 63) source_review, read from each plugin's own .mcp.json. Re-registers this branch's evidence files in manifests/evidence.json against main's copy per the docs/lanes.md hot-file protocol (merged origin/main first) and regenerates the component matrix and grand list. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
…ifest, not the Skills section Item 1's "replace the dangling Skills pointer" sub-bullet pointed the rollback's skills-binary reference at the README's own Skills section, but that section names the binary through its own --skills-bin <ecosystem skills binary> placeholder rather than a concrete path or version -- a reader following the link would land on another placeholder. Cites adoption/skills/manifest.json's cli block directly instead (version "1.7.0"), which is the actual pin. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
The removal table (line 61) sits above "## Codex"; the receipts-table pointer at line 468 correctly says "above" and is unchanged. README re-registered in manifests/evidence.json; validate.py and host_receipts.py validate pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
|
Ready for independent review at head History: an Opus review found 5 findings, then two fix rounds followed, then an Opus fix-verification pass.
|
|
GPT-6 cross-family review at VERDICT: needs_changes
Verified
🤖 Generated with Claude Code |
- Skills rollback: scope every `skills remove` to `-a claude-code codex` (upstream skills@1.7.0 src/remove.ts otherwise targets every known agent format, ~30 of them, not only the two Stage 1 touched); exclude `typesafe-ai` from the automated list, since the pre-Stage-1 backup shows it already linked under claude-code before Stage 1 ran. - Plugin rollback: add `--keep-data` to all three `claude plugin uninstall` commands and document what each retains vs. removes; document that `codex plugin remove` has no equivalent flag and removes the plugin's cache unconditionally. - Workflow/agent/hook/RTK.md rollback: gate every removal on a recorded sha256 digest (repository digest where this document already established byte-identity; a live digest recorded today, labeled as such, otherwise), keeping the file and printing why when it does not match. Pull `contract.config.json` out of automatic removal (no install-time digest exists for a per-host render) and list it as a manual step instead. - Record a superseding claude-code `-4` host receipt: the positive control now requires, and prints only a constant sanitized target for, a `Read` tool_use whose own `input.file_path` (compared by basename only) names the fixture, instead of merely scanning tool_use names. Rehearsed by hand before recording. Updates every README reference to the claude-code receipt chain accordingly. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
…ent-layer-20260927
…ging origin/main origin/main moved by 2 commits touching manifests/evidence.json since this branch's last push. Per docs/lanes.md, took main's copy rather than trust git's clean auto-merge, and re-registered this branch's 9 evidence files (8 host receipts + the Stage 1 client-layer README) with host_receipts.register_file, then re-ran component_matrix.py --write and new_host_grand_list.py --write (both no-ops here; content already matched). Also corrects one README aside that origin/main's merge made stale: 3 of the 10 catalog agent files it lists (isolated-builder, source-scout, stack-verifier) were revised on main after this host's Stage 1 install, so they no longer match the repository copy today. The sha256 gate itself is unaffected (it always compared against a live-read digest, never the repository's), only the informational "matches the repo copy" aside needed correcting. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
…in rollback prose - Reworded the skills -a validation claim: this session read the CLI's agents object keys directly from cli.mjs rather than triggering a live "Invalid agents" error (the extracted 1.7.0 package has no node_modules of its own to run standalone in this environment). - Fixed the Codex context-mode plugin's cache path: it resolves through the context-mode marketplace (~/.codex/plugins/cache/context-mode/context-mode/), not openai-bundled -- this document's own codex -3 receipt excerpt already retains a command under the correct path, so the two now agree. - Reworded the canonical-skill-removal claim to state the CLI's mechanism (keeps the canonical copy and lock entry only if another installed agent still links them, otherwise removes both) rather than assert an outcome for this host that detectInstalledAgents() was never actually run to confirm. - Dropped a dangling "(below)" with nothing to point to. - Confirmed (ls, names only) that the pre-Stage-1 backup's claude/skills/typesafe-ai/ does hold LICENSE and SKILL.md, matching what the README already claimed from the earlier live-symlink check. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
|
Fix round 3 for the GPT-6 cross-family review at
Hot-file protocol: Self-review correction (before this comment): re-reading my own README prose against what I'd actually verified turned up three claims that outran their evidence — fixed in a follow-up commit before this one: (1) the skills Not fixed, flagged as follow-ups, not in this PR's diff:
Checks, all exit 0 at 🤖 Generated with Claude Code |
|
GPT-6 cross-family re-check, round 3, at All four round-2 findings are fixed:
Checks: Residuals. These are nonblocking and recorded under the bounded-loop rule, with no round 4:
🤖 Generated with Claude Code |
|
Merged as
Residuals from GPT-6's round-3 review, recorded here as nonblocking. The bounded-loop rule allows no round 4.
Follow-up tracked separately: #405, fixed by #407 ( |
…squash #391's receipts cited commits published only on its head branch; the repository deletes merged head branches and CI fetches heads and tags only, so main went red until the revisions were tagged. The 2026-09-25 row now says to tag every such revision (receipt-revision/<sha8>) before merging and to check it in a full main-plus-tags clone. This PR's own receipt revision is tagged (receipt-revision/b16b84b8). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
manifests/evidence.json and docs/component-evidence-matrix.md taken from main, this branch's files re-registered, component_matrix and new_host_grand_list regenerated. validate.py passed (7352 hashed files); host_receipts.py validate 192 passed; both --check runs exit 0. Receipt revisions not on main are tagged receipt-revision/<sha8>. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
Date the 2026-09-27 Stage 1 narrative (#391) and cite its coverage section at README.md:414-433 instead of calling it later than the 2026-09-29 check; quote the -76 comment's own words for item 3 instead of "exploratory"; cite AGENTS.md:12 for supported installation; say the edition's :101-107 limits, not dates, the agentmemory evidence; scope the jCodeMunch gap to the Mac. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ons (#662) * docs: preserve PR 508 token-stack history for retirement Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * chore: register PR 508 retirement record evidence Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * docs: apply the PR 662 review round to the PR 508 retirement record Date the 2026-09-27 Stage 1 narrative (#391) and cite its coverage section at README.md:414-433 instead of calling it later than the 2026-09-29 check; quote the -76 comment's own words for item 3 instead of "exploratory"; cite AGENTS.md:12 for supported installation; say the edition's :101-107 limits, not dates, the agentmemory evidence; scope the jCodeMunch gap to the Mac. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore: re-register the PR 508 retirement record after the review round Take main's manifests/evidence.json and re-register only the edited record (sha256 7227ae70...73a7, 21,068 bytes) under the hot-file protocol; both report generators ran and changed nothing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: call the PR 508 merge trigger dormant, not void, in the retirement record Review thread on #662 (record line 220): the #508-merges branch of the overturn condition at docs/decisions/2026-09-30-task-model-routing.md :213-216 cannot fire while #508 stays closed, but the record keeps #508's branch and head, and a closed pull request can be reopened (GitHub GraphQL reopenPullRequest, docs.github.com/en/graphql/reference/pulls #mutation-reopenpullrequest; live schema: "Reopen a pull request."), so a reopened #508 that merges would fire it. Scope ast-grep's :217-218 admission condition to "while #508 stays closed", since routing :201 also ties a Gate A path to #508 merging as written. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore: re-register the PR 508 retirement record after the dormant-trigger fix Re-register only the edited record on the branch's own manifests/evidence.json (sha256 59f58e14...c70d, 21,398 bytes); main is not merged. scripts/component_matrix.py --write and scripts/new_host_grand_list.py --write ran and changed nothing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Scout <scout@local> Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
Scope
native_provenhost receipts (claude-code, codex — claude-code recorded once then superseded twice,-2backing the version claim with a retained command and-3fixing review findings against the checkers and claim text; codex recorded once then superseded three times, the extra-4generation fixing its positive control's command-capture gate, see "Evidence-class table" and "Host receipts" in the README) that evidence Stage 1 of the client-layer install onmac-coordinator-64gb-20260925, plusevidence/artifacts/mac-stage1-client-layer-20260927/README.mdsummarizing the coordinator's backup/rollback, launchctl continuity, headless read-back, plugin revision check, skills status, token-efficiency coverage, Codex profile, cross-host coordination settings and the ai-memory writes every receipt's model turn makes, and the macOS template findings from that run. The client-layer install itself was performed by the coordinator session on that Mac; this PR is the evidence.1aa97765(this branch has mergedorigin/mainfour times, as review passes landed on both sides —c60f3ec5up to5f3a7c21,5b2153c8up to805991bc, and, in this review-fix pass,4212116cup toec27a300thena4d7f5a7up to8bb52541(3 more commits total, none touching a path this PR owns) — verified withgit log --oneline --merges 0f45c0cb..HEAD. Every merge that touchedmanifests/evidence.jsonfolloweddocs/lanes.md's hot-file protocol rather than trusting git's own auto-merge: tookmain's copy and re-registered only this branch's own files throughhost_receipts.py'sregister_file. Re-fetched immediately before pushing and confirmed 0 commits behind).lane:foundation.evidence/hosts/mac-coordinator-64gb-20260925/**,evidence/artifacts/mac-stage1-client-layer-20260927/**,catalogs/landscape/component-evidence-matrix.json,docs/component-evidence-matrix.md— all foundation-owned perdocs/lanes.md.manifests/evidence.jsonis a shared hot file perdocs/lanes.md; it is touched here only for registration, through the documented--write/register_fileprotocol, which does not by itself make this alane:sharedPR (docs/lanes.mdlines 33-38 and 147-148).git diff --name-only origin/main...HEADfrom the worktree, head17ba4b1b, re-run after this pass's merges and fixes): the same shape as before, plus the new codex-4receipt this pass adds —catalogs/landscape/component-evidence-matrix.json,docs/component-evidence-matrix.md,evidence/artifacts/mac-stage1-client-layer-20260927/README.md, the seven claude-code/codex receipt generations (claude-code.json/-2.json/-3.json; codex.json/-2.json/-3.json/-4.json) underevidence/hosts/mac-coordinator-64gb-20260925/, andmanifests/evidence.json. No file under.claude/agents,examples/claude-native/agentsoradoption/agents/claudeis touched, and neither generated grand-list file (catalogs/landscape/new-host-grand-list.json,docs/new-host-grand-list.md) changed. None oforigin/main's own new commits, across any of the three merges, touch any path this PR owns; the only overlap has ever been the sharedmanifests/evidence.jsonregistry.Host request: #382
SOTA sources
docs/contributing-evidence.md(the receipt flow, evidence classes and hot-file protocol) andscripts/host_receipts.py/adoption/host-receipt.schema.json, at this branch's base commit1aa97765.docs/decisions/2026-09-27-mac-single-writer-staged.md(Stage 1 = the client layer, no service change) and host request #382.openai/codexcodex-rs/app-server-daemon/README.mdatrust-v0.157.1, andevidence/receipts/codex-01571-qualification-20260926.json.Evidence-class table
tool_useevents; a separateclaude mcp list(own exit code, no shell pipe) confirms 5 servers all Connectednative_provenevidence/hosts/mac-coordinator-64gb-20260925/mac-coordinator-64gb-20260925--claude-code--use--20260927-3.jsoncodex exec --sandbox read-only --ephemeralturn whose positive control is gated on, and prints, acommand_executionevent whose own command text contains the fixture's filename — not merely the first successfulcommand_executionof any kind (the superseded-3generation's own excerpt shows that weaker gate accepted an unrelated plugin file read; see below) — checked against a negative control;codex mcp list --jsonconfirms config.toml's 8 declared servers are all live and discloses 2 live-only servers (cua_repl,codex_app) without failing on them, run withset -o pipefailso codex's own exit code is never discarded behind the parser's;codex features listcross-checksdaemon_auto_start/hookslive against config.toml, with the same pipefail fix. Does not claimapproval_policy/sandbox_mode, and does not call--sandbox read-onlya safety boundary (see README's "Codex" and "Memory-store writes during recording" sections)native_provenevidence/hosts/mac-coordinator-64gb-20260925/mac-coordinator-64gb-20260925--codex--use--20260927-4.json-2claude-code/codex receipts, and the codex-3receipt: superseded. The-2generations backed the version claim with a retained command but their claim text asserted checks (a Read-tool use,claude mcp list's real pass/fail,codex's exact model command,approval_policy/sandbox_mode) that the commands did not actually perform; the codex-3generation fixed the version-check and disclosure issues but its own positive control still accepted the first successfulcommand_executionof any kind (not one referencing the fixture) and itsmcp list/features listcommands piped codex's own stderr into the parse and discarded codex's own exit code behind the parser's; the claude-code-3and codex-4generations fix the checkers and narrow the claims to match (see the README's "Host receipts" section)native_proven, each superseded...--claude-code--use--20260927.json,...--claude-code--use--20260927-2.json,...--codex--use--20260927.json,...--codex--use--20260927-2.json,...--codex--use--20260927-3.jsonlocal.agent-ecosystem.ai-memory/ollama/qdrantstill running andmaintenancestill not running, in this session's own live re-checknative_proven(for this session's own capture only)launchctl list | grep -E 'agent-ecosystem|native-stack'; see below. Notnative_provenfor "same PIDs as the coordinator's before/after": that comparison relies on the coordinator's own captures, relayed and not independently re-verifiable here (relabeled fromnative_provenafter review; see README "Service continuity")source_reviewscripts/host_requests.py, so it is the same source as the backup claim, not an independent one, and it reports a truncated "4…", not 4,516.)source_reviewfor the coordinator's original capture; the same shape was reproducednative_provenby this session's own claude-code receipt (now-3), recorded later — the same Claude Code session as the original capture, not a second observer (see README intro)readback2.jsonl(not committed: carries a session id, uuid and cwd); reproduced in the receipt aboveadoption_status.pytoken-efficiencyclient_wiring.complete: truesource_reviewskills_status.py/adoption_status.pyoutput, not rerun in this sessionapproval_policy=never,sandbox_mode=danger-full-access,features.daemon_auto_start=false,features.hooks=truesource_reviewfor all four (relabeled fromnative_proven; onlydaemon_auto_start/hooksare cross-checked live, by the receipt above —approval_policy/sandbox_modeare a direct config.toml read, and the receipt's own exec turn runs under--sandbox read-only, notdanger-full-access)~/.codex/config.tomldirectly; the codex receipt above cross-checks livecodex features listfor the two features onlycua_replenabled andcodex_appdisabled (name/enabled/transport only)native_provencodex mcp list --jsoncommand; see README "Codex"source_review— themcp listcommand backing the row above prints no provenance field; this session read each plugin's own.mcp.jsondirectlyremoteControlAtStartup: true,crossSessionInbound: "accept", alongside this profile'sdefaultMode: "bypassPermissions"source_review~/.codex/config.toml's sibling~/.claude/settings.jsondirectly; see README "Cross-host coordination" (new section; not in the previous version of this README)context-mode@context-modeahead of the reviewed revision (stats.jsononly, by GitHub compare API),claude-hud@claude-hudandcodex@openai-codexexact matchessource_reviewgitCommitShavalues and the compare API's{status, files}result are now recorded in the table (they were not retained in either earlier check)source_reviewthroughout, including rtk-hook-wiring and the Qdrant URL: both were checked directly in this session, but neither is inside either receipt's own retainedcommands(relabeled fromnative_proven— a residual from the first review-fix pass that this pass corrects; see README "macOS template findings")Local commands run
Re-run after the two merges and the rtk/Qdrant relabel fix, from the worktree at head
7a045f3d:Re-run after this review-fix pass (rewrote the rollback, disclosed the ai-memory writes,
recorded the codex
-4receipt, relabeled the MCP-provenance claim, fixed thesame-session wording, and the cheap nits — see "Review-fix pass, part 3" below), from the
worktree starting at head
7a045f3d:Full-suite note (unchanged from part 2, not re-run this pass — no code this pass touched
is outside the three modules above):
python3 -m unittest -q(no path filter) was run once before the--writecommands above and reportedRan 6562 tests in 580.371s/FAILED (failures=1, skipped=870), with the one failure beingtest_component_matrix.py::test_real_repository_outputs_are_current— expected, since thechecked-in matrix/grand-list files were still one
--writebehind the two new-3receipts at that point. A second full run was started in the background after
component_matrix.py --write/new_host_grand_list.py --writeabove to confirm thatfailure clears; it includes slow model-adjudication test fixtures (
adjudicate: ...casesthat appear to exercise real provider calls) and was still running well past the first
run's 580s when this PR was pushed, so its result was not waited on here. In its place,
the three specific modules that exercise what this PR changed were run directly and pass
cleanly (see above):
test_host_receipts,test_component_matrixandtest_new_host_grand_list, 261 tests, OK. The full suite is also what CI'svalidatejobruns; its result there is authoritative regardless of this local run.
Decision record
docs/decisions/2026-09-27-mac-single-writer-staged.mdHost evidence
python3 scripts/host_receipts.py validatepasses for every new/changed receipt.-3(claude-code) and-4(codex) receipts can adequacy-check points 1-3 ofdocs/contributing-evidence.md's four review points only; point 4 ("bound to the winner") fails by construction for both (both use--allow-unbound-version, and the tested codex is the ChatGPT app's bundled alpha build, which cannot bind to anopenai/codexrelease pin at all). Perdocs/contributing-evidence.md, a review isagreeonly when all four points hold, so a compliant reviewer's overall verdict here isneeds_changes, which withholdsaccepted/anyplatform_statuschange from these receipts (section 5) — this PR does not ask for or expectacceptedon them. Whether aneeds_changesverdict satisfies this PR's own merge gate (Review, below) is for the coordinator/maintainer to decide;contributing-evidence.mdsection 9's merge gate itself only requires a review to be present or requested, notagree. See the README's "Host receipts" section for the route to bound evidence.platform_statuschange is made from a host receipt alone; every-2and later receipt uses--allow-unbound-version(installed builds are newer than the current catalog pins) and do not by themselves change either component'smacos-arm64status (stilluntested).Checklist
permissions: contents: read(or a narrower, explicitly justified addition). — N/A, no workflow file changed.gitleaks dirover both touched evidence directories reports no leaks (see Local commands run).Review
This PR's receipts carry only the recorder's self-review, from this Claude Code session's
own identity: every receipt generation in this PR, from the coordinator's original
recordings through this pass's codex
-4, carries the identicalrecorded_by.identity_sha256(see the README's intro) — they are all the same session'swork, not independent observations of each other. Independent review needs a different
identity:
scripts/host_receipts.py reviewrefuses a review whose identity equals thereceipt's recorder, and a workflow
agent()or subagent this session launches inheritsthis same session's
$CLAUDE_CODE_SESSION_ID, so it cannot supply one — a same-host reviewrun as a subagent of this session would be refused as a self-review, not accepted as
independent. The viable paths are a fresh Claude Code session on this Mac (its own
session id, started separately from this one, not a subagent it launches) reviewing these
receipts, or the workstation's separate GPT-6 cross-family review requested here; merge
waits for at least one of these, per host request #382's acceptance criteria. As stated in
"Host evidence" above, either review can adequacy-check points 1-3 of the independent-review
rule only; point 4 fails by construction for the current
-3(claude-code) and-4(codex) receipts, so a compliant reviewer's overall verdict is expected to be
needs_changes, notagree. These receipts were never going to reachacceptedwithoutfirst installing the pinned official releases (or, for claude-code only, moving the
landscape pin forward — not possible for codex, whose tested build is a ChatGPT-bundled
alpha and not an
openai/codexrelease at all), and this PR does not ask for that; it asksonly that the review adequacy-check the three points it can. Whether an expected
needs_changesstill satisfies "merge waits for independent review" here is for thecoordinator or a maintainer to decide, not asserted by this PR.
Review-fix pass (2026-09-27)
Applied the review findings on the first version of this PR: relabeled two evidence-class
table rows that had no retained backing (
source_reviewinstead ofnative_proven: thelaunchctl PID-continuity comparison to the coordinator's private captures, and the Codex
profile's
approval_policy/sandbox_mode); recorded-3superseding receipts for bothcomponents with checkers that verify what the claim text says (command text capture and a
negative control for codex; a
Read-tool check and a realclaude mcp listexit-code/statusgate for claude-code; a version-check exit-code bug fixed); disclosed the Codex live MCP
surface's two plugin-provided extra servers and the 13 enabled plugins; disclosed the applied
remoteControlAtStartup/crossSessionInboundsettings, their user scope and thebypassPermissionsinteraction; corrected the backup/rollback section to state what is andis not confirmed and to actually undo Stage 1's additions; recorded the plugin-revision
check's installed SHAs and GitHub compare API output; opened
#394 for the two
pre-existing ecosystem agents' tools-allowlist gaps (host state, not in this repository's
diff); and fixed a wording nit (
manifests/evidence.jsonis a shared hot file touched forregistration only, not foundation-owned).
Review-fix pass, part 2 (2026-09-27)
A second pass, re-verifying the first pass's own fixes against the actual files rather than
trusting its commit message: one row's relabeling had not actually landed. The evidence-class
table (and the README's own "macOS template findings" bullets) still read "rtk-hook-wiring and
the Qdrant URL were independently reconfirmed
native_provenin this session" — the exactfragment the original review flagged as unbacked ("no retained command or output for either
exists anywhere in the diff") — even though the first pass's commit message claimed all three
such rows were relabeled. Confirmed against both
-3receipts' fullcommandsarrays: neitherruns an rtk command, and the codex mcp-list command's own filter prints only
name/enabled/transport, never a server address, by design. Fixed in both the README and this
table; corrected the pass-1 summary above from "three" to "two" rows to match what that pass
actually changed. This pass also merged
origin/main's 2 new commits per the hot-fileprotocol (see "Base commit" above) and re-ran every check listed under "Local commands run".
Full list and disposition of every finding from both review lenses, including the ones
resolved as verification gaps rather than defects and confirmed against the files rather than
assumed, is in this PR's review thread (posted alongside this update).
Review-fix pass, part 3 (2026-09-27)
A third re-review (two lenses, both
needs_changes) against the actual files, all findingsconfirmed before fixing:
rollback": the rollback is now
rsync -a/cp -ponly (a one-line form plus a spelled-outone), with no
--deleteand normanywhere, so it can only add or overwrite a file thebackup holds and can never delete a live file. States plainly that credentials (Codex's
auth.json, Claude Code's own native store), transcripts, sessions, history and cacheswere never in the backup and are never touched, in either direction. Adds a table of every
Stage 1 addition the restore does not remove, each with its own native removal command:
the 10 catalog agents, the two guard hooks,
~/.claude/workflows, the three Claudeplugins (
claude plugin uninstall), theserenauser MCP (claude mcp remove ... -s user), the three Codex MCP servers (codex mcp remove), the Codexcontext-modeplugin(
codex plugin remove),~/.codex/RTK.md,~/.codex/app-server-daemon/settings.json,and the skills manifest (
skills remove --all). Notesapply_claude_settings.py's owntimestamped
settings.jsonbackup as a separate, faster route for that one file.section "Memory-store writes during recording": both clients' user-scope hooks fire on
every receipt's real client turn and write a new session/project into the running
ai-memory service, not a synthetic log;
--ephemeraldoes not stop it; nothing cleans itup. Removed "no running service touched" (replaced with the narrower, accurate claim) and
quoted-and-corrected the frozen claude-code
-3receipt's "no mutation evidenced" and thesuperseded codex
-3receipt's "read-only for safety" (both receipt texts are frozen —receipts change only by recording a new generation — so the correction is in the README's
prose, and the new
-4receipt drops the "for safety" framing at the source). Disclosedthat the declared
context-modeMCP server (default_tools_approval_mode = "approve")runs outside
--sandbox read-only, same as Codex's plugin-provided tools.-3receipt's printed command wasn't the fixture read. Rehearsedthe fix once outside the receipt chain, then recorded
mac-coordinator-64gb-20260925--codex--use--20260927-4.json(
--supersedes ...-3): its positive control now gates on, and prints, acommand_executionwhose own text contains the fixture's filename (
wc -l < fixture.txtthis run), notmerely the first successful
command_executionof any kind. Also fixed while re-recording(cheap alongside the required change):
mcp list/features listnow run withset -o pipefailand discard codex's own stderr before the parse instead of merging it in (socodex's own exit code is never discarded behind
python3's, matching the claude-code-3fix from part 1); a JSON parse failure now prints
len(raw)instead of up to 120 rawcharacters; both
event_kindsand per-item-typeitem_kindsare now printed; a newlimitation states the ai-memory write directly in the receipt. Updated the README's "Host
receipts" table and this PR body's evidence-class table (row for the codex claim, plus the
superseded-receipts row) to describe
-4.split in two: the live count (
native_proven, backed bymcp list) and the provenance —"
cua_repl/codex_appare plugin-provided, not from Stage 1" — nowsource_review(adirect read of each plugin's
.mcp.json, not anything themcp listcommand'sname/enabled/transport-only output shows). Same split in the README's "Codex" section.
observers. They are the same Claude Code session — every receipt in this PR, from the
coordinator's original recordings through today's codex
-4, carries the identicalrecorded_by.identity_sha256. Removed every "independently re-ran"/"independentlyreproduced" in the README and this body; states the shared identity once, plainly, in the
README's intro. Consequence: the "Opus same-host review" this PR's Review section
described would be refused by
scripts/host_receipts.py reviewif run as a subagent ofthis session (same
$CLAUDE_CODE_SESSION_ID, hence the same recorder identity) — rewordedReview and Host evidence to route the same-host independent review to a fresh Claude
Code session on this Mac instead, alongside the workstation's GPT-6 lane.
docs/contributing-evidence.md"point 3 above" →section 8 point 4); corrected the superseded codex
-2receipt's description (thesocraticoderow and the summary line were cut off, notcua_repl/codex_app, whichwere inside the 400-char excerpt); corrected "moving the landscape pins forward" — that
route exists for claude-code but not for codex, whose tested build is a ChatGPT-bundled
alpha and not an
openai/codexrelease at all; corrected the "Base commit" bullet's mergecount from "twice" to three, verified with
git log --oneline --merges; narrowed theclaude-code
-3receipt claim's description in the README to "aReadtool_use ispresent" rather than "reads the file with its own Read tool" (the checker confirms the
former, not the latter).
origin/main's new commits (two more since part 2 landed,ec27a300and1c32ad22, then one more,8bb52541, discovered while re-verifying before push) per thehot-file protocol, and re-ran every check listed under "Local commands run" above.
🤖 Generated with Claude Code
https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu