Skip to content

Stage 1 client-layer evidence: claude-code and codex receipts on mac-coordinator-64gb-20260925 (#382) - #391

Merged
seathatflowsinourveins merged 19 commits into
mainfrom
claude/mac-stage1-client-layer-20260927
Sep 27, 2026
Merged

seathatflowsinourveins merged 19 commits into
mainfrom
claude/mac-stage1-client-layer-20260927

Conversation

@seathatflowsinourveins

@seathatflowsinourveins seathatflowsinourveins commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Scope

  • What this PR changes: adds the two native_proven host receipts (claude-code, codex — claude-code recorded once then superseded twice, -2 backing the version claim with a retained command and -3 fixing review findings against the checkers and claim text; codex recorded once then superseded three times, the extra -4 generation fixing its positive control's command-capture gate, see "Evidence-class table" and "Host receipts" in the README) that evidence Stage 1 of the client-layer install on mac-coordinator-64gb-20260925, plus evidence/artifacts/mac-stage1-client-layer-20260927/README.md summarizing the coordinator's backup/rollback, launchctl continuity, headless read-back, plugin revision check, skills status, token-efficiency coverage, Codex profile, cross-host coordination settings and the ai-memory writes every receipt's model turn makes, and the macOS template findings from that run. The client-layer install itself was performed by the coordinator session on that Mac; this PR is the evidence.
  • Base commit: 1aa97765 (this branch has merged origin/main four times, as review passes landed on both sides — c60f3ec5 up to 5f3a7c21, 5b2153c8 up to 805991bc, and, in this review-fix pass, 4212116c up to ec27a300 then a4d7f5a7 up to 8bb52541 (3 more commits total, none touching a path this PR owns) — verified with git log --oneline --merges 0f45c0cb..HEAD. Every merge that touched manifests/evidence.json followed docs/lanes.md's hot-file protocol rather than trusting git's own auto-merge: took main's copy and re-registered only this branch's own files through host_receipts.py's register_file. Re-fetched immediately before pushing and confirmed 0 commits behind).
  • Lane: lane:foundation.
  • Owned paths touched: evidence/hosts/mac-coordinator-64gb-20260925/**, evidence/artifacts/mac-stage1-client-layer-20260927/**, catalogs/landscape/component-evidence-matrix.json, docs/component-evidence-matrix.md — all foundation-owned per docs/lanes.md. manifests/evidence.json is a shared hot file per docs/lanes.md; it is touched here only for registration, through the documented --write/register_file protocol, which does not by itself make this a lane:shared PR (docs/lanes.md lines 33-38 and 147-148).
  • Verified file list (git diff --name-only origin/main...HEAD from the worktree, head 17ba4b1b, re-run after this pass's merges and fixes): the same shape as before, plus the new codex -4 receipt this pass adds — catalogs/landscape/component-evidence-matrix.json, docs/component-evidence-matrix.md, evidence/artifacts/mac-stage1-client-layer-20260927/README.md, the seven claude-code/codex receipt generations (claude-code .json/-2.json/-3.json; codex .json/-2.json/-3.json/-4.json) under evidence/hosts/mac-coordinator-64gb-20260925/, and manifests/evidence.json. No file under .claude/agents, examples/claude-native/agents or adoption/agents/claude is touched, and neither generated grand-list file (catalogs/landscape/new-host-grand-list.json, docs/new-host-grand-list.md) changed. None of origin/main's own new commits, across any of the three merges, touch any path this PR owns; the only overlap has ever been the shared manifests/evidence.json registry.

Host request: #382

SOTA sources

Evidence-class table

Claim Evidence class Command / receipt
Installed claude-code (2.1.283) does a real headless turn from a scratch cwd, correctly answers a random-fixture line-count positive control, is checked against a deliberately-wrong negative control, and is confirmed to use its Read tool via the transcript's own tool_use events; a separate claude mcp list (own exit code, no shell pipe) confirms 5 servers all Connected native_proven evidence/hosts/mac-coordinator-64gb-20260925/mac-coordinator-64gb-20260925--claude-code--use--20260927-3.json
Installed codex (codex-cli 0.158.0-alpha.2.1, routed through the ChatGPT desktop app's bundled build) does a real codex exec --sandbox read-only --ephemeral turn whose positive control is gated on, and prints, a command_execution event whose own command text contains the fixture's filename — not merely the first successful command_execution of any kind (the superseded -3 generation's own excerpt shows that weaker gate accepted an unrelated plugin file read; see below) — checked against a negative control; codex mcp list --json confirms config.toml's 8 declared servers are all live and discloses 2 live-only servers (cua_repl, codex_app) without failing on them, run with set -o pipefail so codex's own exit code is never discarded behind the parser's; codex features list cross-checks daemon_auto_start/hooks live against config.toml, with the same pipefail fix. Does not claim approval_policy/sandbox_mode, and does not call --sandbox read-only a safety boundary (see README's "Codex" and "Memory-store writes during recording" sections) native_proven evidence/hosts/mac-coordinator-64gb-20260925/mac-coordinator-64gb-20260925--codex--use--20260927-4.json
Original and -2 claude-code/codex receipts, and the codex -3 receipt: superseded. The -2 generations backed the version claim with a retained command but their claim text asserted checks (a Read-tool use, claude mcp list's real pass/fail, codex's exact model command, approval_policy/sandbox_mode) that the commands did not actually perform; the codex -3 generation fixed the version-check and disclosure issues but its own positive control still accepted the first successful command_execution of any kind (not one referencing the fixture) and its mcp list/features list commands piped codex's own stderr into the parse and discarded codex's own exit code behind the parser's; the claude-code -3 and codex -4 generations fix the checkers and narrow the claims to match (see the README's "Host receipts" section) native_proven, each superseded ...--claude-code--use--20260927.json, ...--claude-code--use--20260927-2.json, ...--codex--use--20260927.json, ...--codex--use--20260927-2.json, ...--codex--use--20260927-3.json
local.agent-ecosystem.ai-memory/ollama/qdrant still running and maintenance still not running, in this session's own live re-check native_proven (for this session's own capture only) launchctl list | grep -E 'agent-ecosystem|native-stack'; see below. Not native_proven for "same PIDs as the coordinator's before/after": that comparison relies on the coordinator's own captures, relayed and not independently re-verifiable here (relabeled from native_proven after review; see README "Service continuity")
Backup taken before any write: 4,516 files, a private manifest sha256 source_review coordinator's retained backup manifest, not re-hashed in this session. (The "independently corroborated by #382's status comment" sentence in the previous version of this table is removed: that comment is written by the same coordinator session through scripts/host_requests.py, so it is the same source as the backup claim, not an independent one, and it reports a truncated "4…", not 4,516.)
Headless read-back shape: 18 agents / 67 skills / 11 plugins / 5 MCP servers (all connected) / 107 slash commands / 118 tools source_review for the coordinator's original capture; the same shape was reproduced native_proven by this session's own claude-code receipt (now -3), recorded later — the same Claude Code session as the original capture, not a second observer (see README intro) coordinator's readback2.jsonl (not committed: carries a session id, uuid and cwd); reproduced in the receipt above
Skills status ok 28/28; adoption_status.py token-efficiency client_wiring.complete: true source_review coordinator's retained skills_status.py/adoption_status.py output, not rerun in this session
Codex profile: approval_policy=never, sandbox_mode=danger-full-access, features.daemon_auto_start=false, features.hooks=true source_review for all four (relabeled from native_proven; only daemon_auto_start/hooks are cross-checked live, by the receipt above — approval_policy/sandbox_mode are a direct config.toml read, and the receipt's own exec turn runs under --sandbox read-only, not danger-full-access) this session read ~/.codex/config.toml directly; the codex receipt above cross-checks live codex features list for the two features only
Codex MCP surface: config.toml declares 8 servers; the live surface has 10, cua_repl enabled and codex_app disabled (name/enabled/transport only) native_proven codex receipt above, codex mcp list --json command; see README "Codex"
Codex MCP surface, provenance: both extra servers are plugin-provided, not from Stage 1 source_review — the mcp list command backing the row above prints no provenance field; this session read each plugin's own .mcp.json directly this session's direct file read; see README "Codex"
Cross-host settings applied at user scope: remoteControlAtStartup: true, crossSessionInbound: "accept", alongside this profile's defaultMode: "bypassPermissions" source_review this session read ~/.codex/config.toml's sibling ~/.claude/settings.json directly; see README "Cross-host coordination" (new section; not in the previous version of this README)
Plugin revision check: context-mode@context-mode ahead of the reviewed revision (stats.json only, by GitHub compare API), claude-hud@claude-hud and codex@openai-codex exact matches source_review README "Plugin revision check"; installed gitCommitSha values and the compare API's {status, files} result are now recorded in the table (they were not retained in either earlier check)
macOS template findings (ai-memory 2.4.1 hooks not merged, no Homebrew, no OTel collector, Qdrant port, rtk hook/version, ecosystem-agent tool grants, etc.) source_review throughout, including rtk-hook-wiring and the Qdrant URL: both were checked directly in this session, but neither is inside either receipt's own retained commands (relabeled from native_proven — a residual from the first review-fix pass that this pass corrects; see README "macOS template findings") README "macOS template findings"

Local commands run

$ python3 scripts/host_receipts.py record ... --supersedes ...--claude-code--use--20260927-2 ...   (and the codex equivalent)
evidence/hosts/mac-coordinator-64gb-20260925/mac-coordinator-64gb-20260925--claude-code--use--20260927-3.json
evidence/hosts/mac-coordinator-64gb-20260925/mac-coordinator-64gb-20260925--codex--use--20260927-3.json
exit 0 (both)

$ python3 scripts/host_receipts.py validate
{"receipts": 173, "status": "passed"}
exit 0

$ python3 scripts/component_matrix.py --write   (then --check)
{"flip_rule_violations": 0, "rows": 32, "status": "written"}
{"rows": 32, "status": "checked"}
exit 0

$ python3 scripts/new_host_grand_list.py --write   (then --check)
{"status": "written", "layers": 32, "winners": 66}
{"status": "passed", "layers": 32, "winners": 66}
exit 0

$ python3 scripts/validate.py
{"components": 69, "hashed_files": 7257, "profiles": 4, "receipts": 159, "status": "passed"}
Integrity and scope checks only; no live provider or GPU execution.
exit 0

$ git diff --check
exit 0

$ python3 -m unittest tests.test_host_receipts tests.test_component_matrix tests.test_new_host_grand_list -q
Ran 261 tests in ~11s — OK
exit 0

$ git commit ...   (pre-commit hook runs gitleaks automatically)
0 commits scanned / no leaks found
exit 0

$ gitleaks dir evidence/hosts/mac-coordinator-64gb-20260925 --config .gitleaks.toml --no-banner --redact
no leaks found
exit 0

$ gitleaks dir evidence/artifacts/mac-stage1-client-layer-20260927 --config .gitleaks.toml --no-banner --redact
no leaks found
exit 0

$ launchctl list | grep -E 'agent-ecosystem|native-stack'
(this session, live; PIDs are process ids, not personal data)
local.agent-ecosystem.ai-memory      1879  0
local.agent-ecosystem.maintenance    -     1
local.agent-ecosystem.ollama         1883  0
local.agent-ecosystem.qdrant         1873  0
exit 0

Re-run after the two merges and the rtk/Qdrant relabel fix, from the worktree at head 7a045f3d:

$ git merge origin/main --no-edit   (auto-merged manifests/evidence.json cleanly; corrected below per protocol anyway)
$ git checkout origin/main -- manifests/evidence.json
$ python3 -c '...register_file(...)' <6 receipts + README>   (register_file, once per file)
registered (7 files)

$ python3 scripts/component_matrix.py --write   (then again after the README edit)
{"flip_rule_violations": 0, "rows": 32, "status": "written"}   (both times)

$ python3 scripts/new_host_grand_list.py --write   (then again after the README edit)
{"status": "written", "layers": 32, "winners": 66}   (both times)

$ python3 scripts/host_receipts.py validate
{"receipts": 173, "status": "passed"}
exit 0

$ python3 scripts/validate.py
{"components": 69, "hashed_files": 7261, "profiles": 4, "receipts": 159, "status": "passed"}
exit 0

$ git diff --check
exit 0

$ gitleaks dir evidence/hosts/mac-coordinator-64gb-20260925 --config .gitleaks.toml --no-banner --redact
no leaks found
exit 0

$ gitleaks dir evidence/artifacts/mac-stage1-client-layer-20260927 --config .gitleaks.toml --no-banner --redact
no leaks found
exit 0

$ python3 -m unittest tests.test_host_receipts tests.test_component_matrix tests.test_new_host_grand_list -q
Ran 261 tests in ~11s — OK
exit 0

$ git push   (no force; fast-forward, b521b798..7a045f3d)
exit 0

Re-run after this review-fix pass (rewrote the rollback, disclosed the ai-memory writes,
recorded the codex -4 receipt, relabeled the MCP-provenance claim, fixed the
same-session wording, and the cheap nits — see "Review-fix pass, part 3" below), from the
worktree starting at head 7a045f3d:

$ git merge origin/main --no-edit   (twice, as origin/main moved twice during this pass;
  both auto-merged manifests/evidence.json cleanly, corrected below per protocol anyway)
$ git checkout origin/main -- manifests/evidence.json
$ python3 -c '...register_file(...)' <7 receipts + README>   (register_file, once per file)
registered (8 files)

$ python3 scripts/host_receipts.py record ...   (the new codex -4 receipt, --supersedes ...--codex--use--20260927-3)
evidence/hosts/mac-coordinator-64gb-20260925/mac-coordinator-64gb-20260925--codex--use--20260927-4.json
exit 0

$ python3 scripts/host_receipts.py validate
{"receipts": 174, "status": "passed"}
exit 0

$ python3 scripts/component_matrix.py --write   (then --check)
{"flip_rule_violations": 0, "rows": 32, "status": "written"}
{"rows": 32, "status": "checked"}
exit 0 (both)

$ python3 scripts/new_host_grand_list.py --write   (then --check)
{"status": "written", "layers": 32, "winners": 66}
{"status": "passed", "layers": 32, "winners": 66}
exit 0 (both; neither generated grand-list file's content actually changed)

$ python3 scripts/validate.py
{"components": 69, "hashed_files": 7328, "profiles": 4, "receipts": 159, "status": "passed"}
exit 0

$ git diff --check   (and --cached --check)
exit 0 (both)

$ gitleaks dir evidence/hosts/mac-coordinator-64gb-20260925 --config .gitleaks.toml --no-banner --redact
no leaks found
exit 0

$ gitleaks dir evidence/artifacts/mac-stage1-client-layer-20260927 --config .gitleaks.toml --no-banner --redact
no leaks found
exit 0

$ python3 -m unittest tests.test_host_receipts tests.test_component_matrix tests.test_new_host_grand_list -q
Ran 261 tests in 14.941s — OK
exit 0

$ git commit ...   (pre-commit hook runs gitleaks automatically)
exit 0

$ git push   (no force; fast-forward, 7a045f3d..17ba4b1b)
exit 0

Full-suite note (unchanged from part 2, not re-run this pass — no code this pass touched
is outside the three modules above): python3 -m unittest -q (no path filter) was run once before the
--write commands above and reported Ran 6562 tests in 580.371s /
FAILED (failures=1, skipped=870), with the one failure being
test_component_matrix.py::test_real_repository_outputs_are_current — expected, since the
checked-in matrix/grand-list files were still one --write behind the two new -3
receipts at that point. A second full run was started in the background after
component_matrix.py --write/new_host_grand_list.py --write above to confirm that
failure clears; it includes slow model-adjudication test fixtures (adjudicate: ... cases
that appear to exercise real provider calls) and was still running well past the first
run's 580s when this PR was pushed, so its result was not waited on here. In its place,
the three specific modules that exercise what this PR changed were run directly and pass
cleanly (see above): test_host_receipts, test_component_matrix and
test_new_host_grand_list, 261 tests, OK. The full suite is also what CI's validate job
runs; its result there is authoritative regardless of this local run.

Decision record

docs/decisions/2026-09-27-mac-single-writer-staged.md

Host evidence

  • python3 scripts/host_receipts.py validate passes for every new/changed receipt.
  • Independent review is explicitly requested in this PR (see Review, below): requested from a session with a different recorder identity than this PR's own — a fresh Claude Code session on this Mac, or the workstation's separate GPT-6 cross-family lane (see Review for why a subagent this session launches cannot be that reviewer). A compliant review of the current -3 (claude-code) and -4 (codex) receipts can adequacy-check points 1-3 of docs/contributing-evidence.md's four review points only; point 4 ("bound to the winner") fails by construction for both (both use --allow-unbound-version, and the tested codex is the ChatGPT app's bundled alpha build, which cannot bind to an openai/codex release pin at all). Per docs/contributing-evidence.md, a review is agree only when all four points hold, so a compliant reviewer's overall verdict here is needs_changes, which withholds accepted/any platform_status change from these receipts (section 5) — this PR does not ask for or expect accepted on them. Whether a needs_changes verdict satisfies this PR's own merge gate (Review, below) is for the coordinator/maintainer to decide; contributing-evidence.md section 9's merge gate itself only requires a review to be present or requested, not agree. See the README's "Host receipts" section for the route to bound evidence.
  • No platform_status change is made from a host receipt alone; every -2 and later receipt uses --allow-unbound-version (installed builds are newer than the current catalog pins) and do not by themselves change either component's macos-arm64 status (still untested).

Checklist

  • New/changed GitHub Actions are pinned to a full commit SHA with a version comment (no floating tags). — N/A, no workflow file changed.
  • New/changed workflows declare top-level permissions: contents: read (or a narrower, explicitly justified addition). — N/A, no workflow file changed.
  • No secrets are printed, logged or committed; no new required secret was added without a documented owner. gitleaks dir over both touched evidence directories reports no leaks (see Local commands run).
  • No new paid hosting, subscription or billing surface was introduced.
  • Peer-owned untracked files and worktrees were preserved (not deleted, moved or overwritten).

Review

This PR's receipts carry only the recorder's self-review, from this Claude Code session's
own identity: every receipt generation in this PR, from the coordinator's original
recordings through this pass's codex -4, carries the identical
recorded_by.identity_sha256 (see the README's intro) — they are all the same session's
work, not independent observations of each other. Independent review needs a different
identity: scripts/host_receipts.py review refuses a review whose identity equals the
receipt's recorder, and a workflow agent() or subagent this session launches inherits
this same session's $CLAUDE_CODE_SESSION_ID, so it cannot supply one — a same-host review
run as a subagent of this session would be refused as a self-review, not accepted as
independent. The viable paths are a fresh Claude Code session on this Mac (its own
session id, started separately from this one, not a subagent it launches) reviewing these
receipts, or the workstation's separate GPT-6 cross-family review requested here; merge
waits for at least one of these, per host request #382's acceptance criteria. As stated in
"Host evidence" above, either review can adequacy-check points 1-3 of the independent-review
rule only; point 4 fails by construction for the current -3 (claude-code) and -4
(codex) receipts, so a compliant reviewer's overall verdict is expected to be
needs_changes, not agree. These receipts were never going to reach accepted without
first installing the pinned official releases (or, for claude-code only, moving the
landscape pin forward — not possible for codex, whose tested build is a ChatGPT-bundled
alpha and not an openai/codex release at all), and this PR does not ask for that; it asks
only that the review adequacy-check the three points it can. Whether an expected
needs_changes still satisfies "merge waits for independent review" here is for the
coordinator or a maintainer to decide, not asserted by this PR.

Review-fix pass (2026-09-27)

Applied the review findings on the first version of this PR: relabeled two evidence-class
table rows that had no retained backing (source_review instead of native_proven: the
launchctl PID-continuity comparison to the coordinator's private captures, and the Codex
profile's approval_policy/sandbox_mode); recorded -3 superseding receipts for both
components with checkers that verify what the claim text says (command text capture and a
negative control for codex; a Read-tool check and a real claude mcp list exit-code/status
gate for claude-code; a version-check exit-code bug fixed); disclosed the Codex live MCP
surface's two plugin-provided extra servers and the 13 enabled plugins; disclosed the applied
remoteControlAtStartup/crossSessionInbound settings, their user scope and the
bypassPermissions interaction; corrected the backup/rollback section to state what is and
is not confirmed and to actually undo Stage 1's additions; recorded the plugin-revision
check's installed SHAs and GitHub compare API output; opened
#394 for the two
pre-existing ecosystem agents' tools-allowlist gaps (host state, not in this repository's
diff); and fixed a wording nit (manifests/evidence.json is a shared hot file touched for
registration only, not foundation-owned).

Review-fix pass, part 2 (2026-09-27)

A second pass, re-verifying the first pass's own fixes against the actual files rather than
trusting its commit message: one row's relabeling had not actually landed. The evidence-class
table (and the README's own "macOS template findings" bullets) still read "rtk-hook-wiring and
the Qdrant URL were independently reconfirmed native_proven in this session" — the exact
fragment the original review flagged as unbacked ("no retained command or output for either
exists anywhere in the diff") — even though the first pass's commit message claimed all three
such rows were relabeled. Confirmed against both -3 receipts' full commands arrays: neither
runs an rtk command, and the codex mcp-list command's own filter prints only
name/enabled/transport, never a server address, by design. Fixed in both the README and this
table; corrected the pass-1 summary above from "three" to "two" rows to match what that pass
actually changed. This pass also merged origin/main's 2 new commits per the hot-file
protocol (see "Base commit" above) and re-ran every check listed under "Local commands run".
Full list and disposition of every finding from both review lenses, including the ones
resolved as verification gaps rather than defects and confirmed against the files rather than
assumed, is in this PR's review thread (posted alongside this update).

Review-fix pass, part 3 (2026-09-27)

A third re-review (two lenses, both needs_changes) against the actual files, all findings
confirmed before fixing:

  • [blocking] Rollback could delete data the backup never held. Rewrote "Backup and
    rollback": the rollback is now rsync -a/cp -p only (a one-line form plus a spelled-out
    one), with no --delete and no rm anywhere, so it can only add or overwrite a file the
    backup holds and can never delete a live file. States plainly that credentials (Codex's
    auth.json, Claude Code's own native store), transcripts, sessions, history and caches
    were never in the backup and are never touched, in either direction. Adds a table of every
    Stage 1 addition the restore does not remove, each with its own native removal command:
    the 10 catalog agents, the two guard hooks, ~/.claude/workflows, the three Claude
    plugins (claude plugin uninstall), the serena user MCP (claude mcp remove ... -s user), the three Codex MCP servers (codex mcp remove), the Codex context-mode plugin
    (codex plugin remove), ~/.codex/RTK.md, ~/.codex/app-server-daemon/settings.json,
    and the skills manifest (skills remove --all). Notes apply_claude_settings.py's own
    timestamped settings.json backup as a separate, faster route for that one file.
  • [should_fix] Every receipt's model turn wrote to the live ai-memory store. New README
    section "Memory-store writes during recording": both clients' user-scope hooks fire on
    every receipt's real client turn and write a new session/project into the running
    ai-memory service, not a synthetic log; --ephemeral does not stop it; nothing cleans it
    up. Removed "no running service touched" (replaced with the narrower, accurate claim) and
    quoted-and-corrected the frozen claude-code -3 receipt's "no mutation evidenced" and the
    superseded codex -3 receipt's "read-only for safety" (both receipt texts are frozen —
    receipts change only by recording a new generation — so the correction is in the README's
    prose, and the new -4 receipt drops the "for safety" framing at the source). Disclosed
    that the declared context-mode MCP server (default_tools_approval_mode = "approve")
    runs outside --sandbox read-only, same as Codex's plugin-provided tools.
  • [should_fix] The codex -3 receipt's printed command wasn't the fixture read. Rehearsed
    the fix once outside the receipt chain, then recorded mac-coordinator-64gb-20260925--codex--use--20260927-4.json
    (--supersedes ...-3): its positive control now gates on, and prints, a command_execution
    whose own text contains the fixture's filename (wc -l < fixture.txt this run), not
    merely the first successful command_execution of any kind. Also fixed while re-recording
    (cheap alongside the required change): mcp list/features list now run with set -o pipefail and discard codex's own stderr before the parse instead of merging it in (so
    codex's own exit code is never discarded behind python3's, matching the claude-code -3
    fix from part 1); a JSON parse failure now prints len(raw) instead of up to 120 raw
    characters; both event_kinds and per-item-type item_kinds are now printed; a new
    limitation states the ai-memory write directly in the receipt. Updated the README's "Host
    receipts" table and this PR body's evidence-class table (row for the codex claim, plus the
    superseded-receipts row) to describe -4.
  • [should_fix] MCP-surface provenance was labeled stronger than what was checked. Row 30
    split in two: the live count (native_proven, backed by mcp list) and the provenance —
    "cua_repl/codex_app are plugin-provided, not from Stage 1" — now source_review (a
    direct read of each plugin's .mcp.json, not anything the mcp list command's
    name/enabled/transport-only output shows). Same split in the README's "Codex" section.
  • [should_fix] Coordinator session and evidence-PR session treated as independent
    observers.
    They are the same Claude Code session — every receipt in this PR, from the
    coordinator's original recordings through today's codex -4, carries the identical
    recorded_by.identity_sha256. Removed every "independently re-ran"/"independently
    reproduced" in the README and this body; states the shared identity once, plainly, in the
    README's intro. Consequence: the "Opus same-host review" this PR's Review section
    described would be refused by scripts/host_receipts.py review if run as a subagent of
    this session (same $CLAUDE_CODE_SESSION_ID, hence the same recorder identity) — reworded
    Review and Host evidence to route the same-host independent review to a fresh Claude
    Code session on this Mac instead, alongside the workstation's GPT-6 lane.
  • Nits: fixed a wrong cross-reference (docs/contributing-evidence.md "point 3 above" →
    section 8 point 4); corrected the superseded codex -2 receipt's description (the
    socraticode row and the summary line were cut off, not cua_repl/codex_app, which
    were inside the 400-char excerpt); corrected "moving the landscape pins forward" — that
    route exists for claude-code but not for codex, whose tested build is a ChatGPT-bundled
    alpha and not an openai/codex release at all; corrected the "Base commit" bullet's merge
    count from "twice" to three, verified with git log --oneline --merges; narrowed the
    claude-code -3 receipt claim's description in the README to "a Read tool_use is
    present" rather than "reads the file with its own Read tool" (the checker confirms the
    former, not the latter).
  • Also merged origin/main's new commits (two more since part 2 landed, ec27a300 and
    1c32ad22, then one more, 8bb52541, discovered while re-verifying before push) per the
    hot-file protocol, and re-ran every check listed under "Local commands run" above.

🤖 Generated with Claude Code

https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu

…b-20260925 (#382)

Records native_proven use-stage receipts for the installed claude-code and
codex clients on the Mac's Stage 1 client layer (host request #382): a
version-resolving command, a headless claude -p turn with a random-fixture
positive control plus claude mcp list, and a codex exec turn (JSON event
stream, command_execution positive control) plus codex mcp list --json and
codex features list cross-checked against this host's own config.toml. Both
receipts are recorded with --allow-unbound-version (installed builds are
newer than the current catalog pins) and superseded once each to back the
version claim with a retained command instead of only --component-version.
Refreshes the derived component-evidence-matrix accordingly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Merges origin/main (no conflicts) and re-registers the derived matrix/grand-
list files. Adds evidence/artifacts/mac-stage1-client-layer-20260927/README.md:
the coordinator's Stage 1 client-layer install on mac-coordinator-64gb-20260925
(backup summary and rollback, launchctl before/after plus a fresh live re-check,
the headless read-back names/counts, plugin revision check, skills status,
token-efficiency/client-wiring coverage, the Codex profile, and the macOS
template findings), each line marked as coordinator-reported or independently
re-checked in this session. Notes two additional findings this session found
while recording the codex receipt: the installed codex-cli routes through the
ChatGPT desktop app's bundled build, and codex exec needs stdin redirected
from /dev/null or it blocks indefinitely.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@seathatflowsinourveins seathatflowsinourveins added the lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers label Sep 27, 2026
…TH inference

Reads ~/.claude/settings.json's own env.PATH value and re-runs command -v
rtk / rtk --version with exactly that PATH (the environment Claude Code
gives its hooks) instead of inferring resolution order from the PATH list.
Confirms it resolves to the ecosystem install and reports rtk 0.50.0.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…cord receipts with verified checks, disclose Codex/cross-host surface

Applies the round of review findings against PR #391:

- Record -3 generations of both host receipts, superseding -2: fix the codex
  version-check command's exit-code bug (a later echo, not `codex --version`,
  decided it), capture the model's actual command_execution text instead of
  assuming `wc -l`, disclose live-only MCP servers (cua_repl, codex_app)
  without failing the receipt on them, drop the unchecked approval_policy/
  sandbox_mode assertion, add a Read-tool-use check and a real claude mcp
  list exit-code/status gate to the claude-code receipt, and add an
  in-transcript negative control to both.
- Relabel three PR-body evidence-class rows from native_proven to
  source_review where no retained command backs the claim (launchctl
  PID-continuity comparison to the coordinator's private captures, the
  Codex profile's approval_policy/sandbox_mode, and the rtk-hook/Qdrant-URL
  "independently reconfirmed" line), and remove a corroboration claim that
  shared its source with what it claimed to corroborate.
- README: disclose the Codex live MCP surface (10 servers vs. 8 declared in
  config.toml, the two extras plugin-provided) and the applied
  remoteControlAtStartup/crossSessionInbound settings, their user scope and
  interaction with this profile's bypassPermissions default; correct the
  backup/rollback section to state what is and is not confirmed and to
  actually undo Stage 1's additions instead of merging over them; record the
  plugin-revision check's installed SHAs and GitHub compare API output;
  state that a compliant independent review of the -3 receipts cannot
  return agree (both are --allow-unbound-version).
- Open #394 for the two pre-existing ecosystem agents' tools-allowlist gaps
  (host state, not in this repository's diff).
- Fix a lanes.md wording nit: manifests/evidence.json is a shared hot file
  touched for registration only, not foundation-owned.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…re source_review

Both claims were checked directly in this session, not through either host
receipt's own retained `commands` (neither's command array prints an rtk
version or a server address), so they were still labeled `native_proven` in
the README even after the prior review-fix pass (b521b79), which relabeled
three other rows the same way but left this exact fragment ("rtk-hook-wiring
and the Qdrant URL were independently reconfirmed native_proven") unchanged.
Confirmed against both -3 receipts' full command arrays: neither runs an rtk
command, and the codex mcp-list command's own filter only ever prints
name/enabled/transport, never a server URL, by design (to avoid capturing
MCP server config in the receipt).

README: both bullets now state the claim is source_review, matching the
wording already used for the Codex profile and launchctl-continuity claims
in the same document. The PR body's evidence-class table gets the same
relabel via gh pr edit (not a commit), alongside refreshed base-commit/
verified-file-list/head references for the origin/main merge below.

Also completes the hot-file merge this commit's parent started: took main's
manifests/evidence.json (2 new commits landed upstream: 805991b, 6c073fc)
rather than trust the clean auto-merge, and re-registered only this branch's
7 files (6 receipts + README) through host_receipts.py's register_file.

Re-verified: python3 scripts/host_receipts.py validate,
scripts/component_matrix.py --write, scripts/new_host_grand_list.py --write,
scripts/validate.py (FAILS=0), git diff --check, gitleaks over both touched
evidence directories (0 leaks), and the three targeted unittest modules
(tests.test_host_receipts, tests.test_component_matrix,
tests.test_new_host_grand_list -- 261 tests, OK).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Review findings: full disposition

Both saved Opus reviews (review-evidence:382-stage1-client-layer,
review-security:382-stage1-client-layer) read and confirmed against the actual files
before acting, per finding. New head: 7a045f3d. Nothing was refuted — every should_fix
and nit from both lenses held up against the source; disposition below.

review-evidence lens (5 should_fix, 8 nit)

Finding (short) Disposition
Evidence-class table labels 3 claims native_proven with no retained backing (launchctl PID-continuity, Codex approval_policy/sandbox_mode, rtk-hook/Qdrant-URL) 2 of 3 fixed in b521b798. The 3rd (rtk-hook/Qdrant-URL) was not actually relabeled despite that commit's message — confirmed by reading both -3 receipts' full commands arrays: neither prints an rtk version or a server address. Fixed now in 7a045f3d (README + this PR's evidence-class table).
-2 receipts assert checks their commands never ran (assumed wc -l, unchecked approval_policy/sandbox_mode, a version-check exit-code bug, claude mcp list with no real pass/fail gate) Fixed in b521b798: -3 generations recorded with checkers that verify exactly the claim text (real command_execution text capture, an in-transcript Read-tool-use check, a claude mcp list exit-code+status gate, the version check's own exit code). Verified by reading both -3 JSON files directly.
Point 4 ("bound to the winner") can't be met; PR must say a compliant review can't return agree Fixed in b521b798: stated explicitly in the README's "Host receipts" section and the PR body's "Host evidence"/"Review" sections.
Rollback doesn't match backup scope, merges instead of replacing Fixed in b521b798: rsync --delete rollback that undoes Stage 1's additions; explicit "not confirmed" on whether ~/.claude.json is in the backup manifest. Re-checked this session's own retained scratch files (launchctl-*.txt etc.) for anything settling that — nothing does, so the hedge stays as-is rather than being asserted either way.
Plugin-revision check has no installed SHAs / compare-API output Fixed in b521b798: table now has installed gitCommitSha + compare API {status,files} per plugin, and a PR-body evidence-class row.
mcp-list excerpt truncated before summary/live-only-server line Fixed in b521b798: summary line (declared_count/live_count/live_only_not_in_config_toml) now printed first in the -3 receipt, ahead of the 400-char truncation point.
Weak positive control (1-in-6 guess), no negative control Fixed in b521b798: fixture range widened to 7–96 in both -3 receipts; both add an in-transcript negative control, confirmed rejected.
prerequisites_missing explanation wrong (called a per-profile gap; it's manifest-wide) Fixed in b521b798: README now correctly attributes it to adoption/manifest.json's manifest-wide supported_platforms, and states the Python version used.
Codex-exec-needs-/dev/null presented as new, not cited to existing 0.155.1 evidence Fixed in b521b798: labeled an unretained observation and cited gap-wave2-20260923/foundation__quality-evaluation/README.md's prior finding on the same upstream behavior.
PR body says touched paths "all foundation-owned" incl. manifests/evidence.json Fixed in b521b798: PR body now says it's a shared hot file touched only for registration (matches the finding's suggested wording).
Verification gap: git diff --name-only, validate.py/CI on head Supplied: "Verified file list" bullet + "Local commands run" in the PR body, re-run and refreshed at 7a045f3d after both origin/main merges.
Verification gap: gitleaks exit code, full unittest, CI validate/validate-macos Supplied: gitleaks (0 leaks, both dirs) and unittest output in "Local commands run"; CI checks listed in this comment's CI section below.
Verification gap: PR label, #382 block reason, decision requirement, backup/launchctl Supplied as far as read-only evidence allows: PR label confirmed lane:foundation; README's "Limits" quotes #382's exact request:blocked status-comment reason. Backup manifest and launchctl are the coordinator's private/live state — this session doesn't touch services or re-open that, per this task's own constraints, so those two stay marked unverified rather than asserted.

review-security lens (3 should_fix, 5 nit, 1 no-findings)

Finding (short) Disposition
Codex tool surface understated (10 live MCP servers vs. 8 declared; 13 enabled plugins incl. computer-use/browser) Fixed in b521b798: README's "Codex" section discloses cua_repl/codex_app, all 13 enabled plugins, and that the read-only sandbox doesn't restrict plugin tools.
Cross-host settings (remoteControlAtStartup, crossSessionInbound: accept) applied but unrecorded; undercuts security premise under bypassPermissions Fixed in b521b798: new README "Cross-host coordination" section states both settings, their user scope, and the bypassPermissions interaction; framed as disclosure, not endorsement, with the scoping decision left to the owner. Re-confirmed this session by a value-free key check of ~/.claude/settings.json: remoteControlAtStartup=true, crossSessionInbound="accept", defaultMode="bypassPermissions", skipDangerousModePermissionPrompt=true, isolatePeerMachines absent — all match what the README states.
Ecosystem-worker/researcher agents' tool grants not named (Serena symbol-edit tools via no tools: line; wildcard MCP grant) Fixed in b521b798: README names the concrete grants; #394 opened and confirmed to exist with matching content (host-state cleanup, not in this repo's diff).
claude -p with no --permission-mode/--allowedTools (PERM-03) Documented in b521b798: both -3 receipts' limitations disclose this as a known, unremedied deviation; finding's own fix said no re-record needed.
Codex binary is a ChatGPT-app-bundled launcher, can change on app update with no receipt Documented in b521b798: README's Codex section states this and names the remedy (a pinned openai/codex release first on PATH), not remedied by this PR.
gitleaks/secret-scan not evidenced Supplied in b521b798 and re-run at 7a045f3d: gitleaks dir over both touched directories, 0 leaks each time; CI secret-scan in progress at time of this comment (see below).
Changed-file-set not independently confirmed Supplied: "Verified file list" bullet, refreshed at 7a045f3d.
updater.autoUpdateEnabled / shell_environment_policy.set relayed, not re-checked Left correctly unverified in b521b798: README states these stay source_review, not backed by any receipt (the codex receipt proves only features.daemon_auto_start=false, a different setting).
No-findings disposition (privacy/mutation sweep) No action required, as the finding itself states.

What's outstanding (not fixable from this PR alone)

  • Both independent reviews are still outstanding — 0 PR reviews as of this comment. Per docs/contributing-evidence.md, a compliant review of the current -3 receipts can adequacy-check points 1–3 only; point 4 ("bound to the winner") fails by construction (both receipts are --allow-unbound-version; the tested codex is the ChatGPT app's bundled alpha build, which can't bind to an openai/codex release pin at all). This PR does not ask for or expect accepted on these receipts — see README "Host receipts" and PR body "Review".
  • Owner decisions, not resolved by this PR: scope crossSessionInbound to participating sessions vs. accept the user-wide risk under bypassPermissions; whether the Codex plugin-provided computer-use/browser surface is accepted for a never/danger-full-access profile; reconciling issue [mac-coordinator] other: Stage 1 client layer on the 64 GB Mac (no service changes) #382's request:blocked label with the now-working state the receipts show; a pinned official openai/codex release ahead of the ChatGPT-bundled one on PATH.
  • ecosystem-worker/ecosystem-researcher fail the workflow contract's tools-allowlist checks #394 tracks the two ecosystem agents' tools-allowlist gaps (host state, separate from this diff).

Hot-file protocol

origin/main moved by 2 commits (805991bc, 6c073fcc) between the previous push and this
one, with one overlap: manifests/evidence.json. Per docs/lanes.md, took main's copy and
re-registered only this branch's 7 files (6 receipts + README) rather than trust git's clean
auto-merge. Re-ran host_receipts.py validate, both --write/--check pairs, validate.py
(FAILS=0), and git diff --check after.

CI (at time of this comment)

18 checks passing, 3 skipped (expected), 3 in progress: validate, validate-macos,
secret-scan. No failures.

…ry writes, record codex -4 with a fixture-gated positive control

Rewrites the README rollback so it can only add or overwrite files the
private backup holds (rsync/cp, no --delete, no rm) and lists every Stage 1
addition it does not remove with its own native removal command. Discloses
that every receipt's model turn fires the live ai-memory hooks and writes to
the running production store, and that the declared context-mode MCP server
and Codex's plugin-provided tools both run outside --sandbox read-only.
Records mac-coordinator-64gb-20260925--codex--use--20260927-4.json
(supersedes -3): its positive control now gates on and prints a
command_execution whose command text actually references the fixture,
rather than the first successful command_execution of any kind; also fixes
mcp list/features list to use set -o pipefail with stderr discarded before
the parse. Relabels the MCP-surface provenance sub-claim source_review
instead of native_proven in the README and PR body. Fixes wording that
treated the coordinator session and this evidence-PR session as independent
observers -- every receipt in this PR carries the identical
recorded_by.identity_sha256, so they are the same session -- and reroutes
the same-host independent review to a fresh session accordingly. Plus the
cross-reference, excerpt-description, pin-wording and merge-count nits from
the re-review.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Review-fix pass, part 3: addressed both lenses' re-review (full disposition in the PR body's new "Review-fix pass, part 3" section).

  • [blocking] Rewrote the README rollback: rsync -a/cp -p only, no --delete, no rm anywhere — it can only add or overwrite a file the backup holds, never delete a live one. States plainly that credentials, transcripts, sessions, history and caches were never in the backup and are never touched. Lists every Stage 1 addition the restore doesn't remove, each with its own native removal command (agents, guard hooks, workflows, plugins, MCP servers, RTK.md, app-server-daemon settings, skills).
  • [should_fix] Disclosed that every receipt's model turn fires the live ai-memory hooks and writes a new session/project into the running production store (new README section "Memory-store writes during recording"). Corrected "no running service touched", the frozen claude-code -3 receipt's "no mutation evidenced", and the codex -3 receipt's "read-only for safety" (receipts are immutable, so the correction is in the README's prose; the new -4 generation drops the phrase at the source). Disclosed that the declared context-mode MCP server (approved individually) and Codex's plugin-provided tools both run outside --sandbox read-only.
  • [should_fix] Recorded mac-coordinator-64gb-20260925--codex--use--20260927-4.json (supersedes -3): rehearsed the fix outside the receipt chain first, then the real recording — its positive control now gates on and prints a command_execution whose command text actually contains the fixture's filename, not merely the first successful command_execution of any kind (the -3 gate had accepted an unrelated plugin SKILL.md read). Also fixed mcp list/features list to run with set -o pipefail and stderr kept out of the JSON/text parse, and to print len(raw) instead of up to 120 raw characters on a parse failure.
  • [should_fix] Relabeled the MCP-surface provenance sub-claim ("cua_repl/codex_app are plugin-provided, not from Stage 1") source_review, not native_proven, in both the README and the PR body's evidence-class table (row split in two).
  • [should_fix] Fixed wording treating the coordinator session and this evidence-PR session as independent observers — every receipt in this PR carries the identical recorded_by.identity_sha256, so they're the same Claude Code session. Removed "independently re-ran"/"independently reproduced" throughout; rerouted the "Opus same-host review" to a fresh Claude Code session (a subagent of this session would be refused by host_receipts.py review as a self-review).
  • Nits: fixed a wrong docs/contributing-evidence.md cross-reference; corrected the superseded codex -2 receipt's excerpt description; corrected "moving the landscape pins forward" (works for claude-code, not for codex's ChatGPT-bundled build); corrected the base-commit merge count (now 4, verified with git log --oneline --merges).

Checks (all from the worktree at the new head):

  • host_receipts.py validate — {"receipts": 174, "status": "passed"} — exit 0
  • component_matrix.py --write / --check — {"flip_rule_violations": 0, "rows": 32, "status": "written"} / {"rows": 32, "status": "checked"} — exit 0
  • new_host_grand_list.py --write / --check — {"status": "written", "layers": 32, "winners": 66} / {"status": "passed", ...} — exit 0 (neither generated file's content changed)
  • validate.py — {"components": 69, "hashed_files": 7328, "profiles": 4, "receipts": 159, "status": "passed"} — exit 0
  • git diff --check (working tree and --cached) — exit 0
  • gitleaks dir on both touched evidence directories — no leaks found
  • python3 -m unittest tests.test_host_receipts tests.test_component_matrix tests.test_new_host_grand_list -q — 261 tests, OK

New head: 17ba4b1b (was 7a045f3d).

… ordering and receipt citations

Fixes nine defects from an independent Opus re-review of the Stage 1
client-layer README:

- Replaces the unscoped `<skills-bin> remove --all` with one
  `skills remove <name> -g -y` per pinned skill (all 28), citing the
  decision record's -g scoping rule and lifecycle.md's "never uninstall
  all user tools to roll back one package".
- Replaces the directory-wide workflows delete with a name-and-sha256
  verified per-file removal of the 13 files Stage 1 actually placed in
  ~/.claude/workflows (12 byte-identical to the repo source, one a
  per-host template render).
- Reorders the rollback so every removal runs before the restore -- a
  restore run first would overwrite the live state the removal table
  was built from -- and rescopes the "never runs rm" claim to the
  restore step alone, since removal now legitimately runs rm.
- Joins the one-line restore's three commands with `;` instead of `&&`
  so a missing claude.json backup can't silently skip the codex rsync.
- Documents that only two of the four hook files
  install_claude_profile.py's HOOKS map can place actually exist on
  this host; the other two were added to that map after Stage 1's own
  install ran here.
- Replaces the blind rm of app-server-daemon/settings.json with a
  conditional that restores from the private backup if it holds a
  copy, otherwise removes only the one key this session is known to
  have written, since the file's prior existence is unrecorded.
- Corrects the claim that both current receipts print a sanitized
  command line: only codex -4 does; claude-code -3 prints tool names
  only.
- Fixes a section citation (contributing-evidence.md section 3 step 8,
  not section 8).
- Adds a sentence labeling the -4 receipt's plugin-to-server
  attribution (line 63) source_review, read from each plugin's own
  .mcp.json.

Re-registers this branch's evidence files in manifests/evidence.json
against main's copy per the docs/lanes.md hot-file protocol (merged
origin/main first) and regenerates the component matrix and grand
list.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
…ifest, not the Skills section

Item 1's "replace the dangling Skills pointer" sub-bullet pointed the
rollback's skills-binary reference at the README's own Skills section,
but that section names the binary through its own --skills-bin
<ecosystem skills binary> placeholder rather than a concrete path or
version -- a reader following the link would land on another
placeholder. Cites adoption/skills/manifest.json's cli block directly
instead (version "1.7.0"), which is the actual pin.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
The removal table (line 61) sits above "## Codex"; the receipts-table
pointer at line 468 correctly says "above" and is unchanged. README
re-registered in manifests/evidence.json; validate.py and
host_receipts.py validate pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Ready for independent review at head 327d139d99f85a31d9cc50321838fec210278951. GPT-6 cross-family review is requested from the workstation.

History: an Opus review found 5 findings, then two fix rounds followed, then an Opus fix-verification pass.

  • Finding status:
    • Findings 2-5 were fixed in round 1: the ai-memory write disclosure, the codex -4 positive control, the source_review relabel, and the "independently" wording.
    • Finding 1 (rollback) was partly fixed in round 1. Round 2 completed it: per-skill skills remove <name> -g -y for the 28 manifest skills instead of --all, 13 per-file workflow removals checked by sha256 instead of rm -rf, ; instead of && in the one-line rollback, a conditional rollback of the app-server-daemon updater key, and removal before restore.
  • Checks at this head: host_receipts.py validate 174 passed; validate.py passed (7330 hashed files); component_matrix.py --check and new_host_grand_list.py --check exit 0; git diff --check clean.
  • Not independent: every re-run of the receipts was the recording session's own (the identical recorded_by.identity_sha256). The README says so.

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

GPT-6 cross-family review at 327d139d (gpt-6-astra, effort max, read-only; run from the workstation session). The review text follows verbatim.

VERDICT: needs_changes

  • High — README.md:150: Skill removal still reaches other tools. In skills v1.7.0’s removal implementation, omitting -a selects every known agent and recursively removes matching skill directories. These commands can delete a same-named Cursor skill outside the backup. Acceptable fix: remove only recorded Stage 1 additions, restrict agents with -a claude-code codex, and preserve pre-existing/shared skills.

  • High — README.md:57: Plugin uninstall can delete persistent data. All three commands omit --keep-data. Installed Claude 2.1.283 help and official documentation confirm that uninstalling the last installation deletes its persistent data directory. This contradicts the rollback’s preservation promise. Acceptable fix: add --keep-data and accurately distinguish retained data from removed installation files.

  • Medium — README.md:91: Workflow removals have no checksum guards. The 13 commands are unconditional rm calls. A historical comparison does not protect subsequent user edits, and the rendered contract.config.json has no retained installed checksum. Acceptable fix: gate each removal on its recorded Stage 1 digest, preserving mismatches. Apply the same protection to the agent, hook, and RTK file removals.

  • Medium — Claude receipt -3:48: The fixture-read claim remains overstated. The checker inspects tool names, not Read inputs. Its unchanged code accepted a synthetic transcript containing an unrelated Read and a Bash fixture count. README clarification does not correct the current receipt’s claim. Acceptable fix: supersede it with a narrower claim or verify and retain a sanitized fixture-target check.

Verified

  • Receipt validation, scripts/validate.py, both generated-view checks, 261 targeted tests, both gitleaks scans, and whitespace checks passed.
  • Full CI suite at the exact head: 6,624 tests passed, 698 skipped. The local full-suite attempt exceeded the MCP response timeout and was stopped; no local full-suite pass is claimed.
  • Both versions are explicitly off-pin. Matrix statuses remain untested; receipt counts do not promote acceptance.
  • Mac execution remains historical evidence; these local checks establish artifact consistency.
  • No private literals found in the diff. Checkout remains clean; nothing edited, committed, or posted.

🤖 Generated with Claude Code

https://claude.ai/code/session_012qwXNtFyUhY5x2dtrm7tG5

- Skills rollback: scope every `skills remove` to `-a claude-code codex`
  (upstream skills@1.7.0 src/remove.ts otherwise targets every known agent
  format, ~30 of them, not only the two Stage 1 touched); exclude
  `typesafe-ai` from the automated list, since the pre-Stage-1 backup shows
  it already linked under claude-code before Stage 1 ran.
- Plugin rollback: add `--keep-data` to all three `claude plugin uninstall`
  commands and document what each retains vs. removes; document that
  `codex plugin remove` has no equivalent flag and removes the plugin's
  cache unconditionally.
- Workflow/agent/hook/RTK.md rollback: gate every removal on a recorded
  sha256 digest (repository digest where this document already established
  byte-identity; a live digest recorded today, labeled as such, otherwise),
  keeping the file and printing why when it does not match. Pull
  `contract.config.json` out of automatic removal (no install-time digest
  exists for a per-host render) and list it as a manual step instead.
- Record a superseding claude-code `-4` host receipt: the positive control
  now requires, and prints only a constant sanitized target for, a `Read`
  tool_use whose own `input.file_path` (compared by basename only) names
  the fixture, instead of merely scanning tool_use names. Rehearsed by hand
  before recording. Updates every README reference to the claude-code
  receipt chain accordingly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
…ging origin/main

origin/main moved by 2 commits touching manifests/evidence.json since this
branch's last push. Per docs/lanes.md, took main's copy rather than trust
git's clean auto-merge, and re-registered this branch's 9 evidence files
(8 host receipts + the Stage 1 client-layer README) with
host_receipts.register_file, then re-ran component_matrix.py --write and
new_host_grand_list.py --write (both no-ops here; content already matched).

Also corrects one README aside that origin/main's merge made stale: 3 of
the 10 catalog agent files it lists (isolated-builder, source-scout,
stack-verifier) were revised on main after this host's Stage 1 install, so
they no longer match the repository copy today. The sha256 gate itself is
unaffected (it always compared against a live-read digest, never the
repository's), only the informational "matches the repo copy" aside needed
correcting.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
…in rollback prose

- Reworded the skills -a validation claim: this session read the CLI's
  agents object keys directly from cli.mjs rather than triggering a live
  "Invalid agents" error (the extracted 1.7.0 package has no node_modules
  of its own to run standalone in this environment).
- Fixed the Codex context-mode plugin's cache path: it resolves through
  the context-mode marketplace (~/.codex/plugins/cache/context-mode/context-mode/),
  not openai-bundled -- this document's own codex -3 receipt excerpt
  already retains a command under the correct path, so the two now agree.
- Reworded the canonical-skill-removal claim to state the CLI's mechanism
  (keeps the canonical copy and lock entry only if another installed agent
  still links them, otherwise removes both) rather than assert an outcome
  for this host that detectInstalledAgents() was never actually run to
  confirm.
- Dropped a dangling "(below)" with nothing to point to.
- Confirmed (ls, names only) that the pre-Stage-1 backup's
  claude/skills/typesafe-ai/ does hold LICENSE and SKILL.md, matching what
  the README already claimed from the earlier live-symlink check.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

Fix round 3 for the GPT-6 cross-family review at 327d139d. New head: cd0d37eb.

  • High — README:150, skill removal reaches other tools. Added -a claude-code codex to all 27 skills remove lines (order matters: the CLI's own arg parser requires -a last, after -g -y, since it greedily consumes trailing bare words as agent names). Confirmed claude-code/codex are the CLI's own agents object keys by reading cli.mjs directly (not by a live run — the extracted 1.7.0 package has no node_modules to run standalone here). Excluded the 28th skill, typesafe-ai, from the automated list: the pre-Stage-1 backup already holds claude/skills/typesafe-ai/{LICENSE,SKILL.md} (ls, names only, confirmed against the backup path itself, not just the live symlink target), so its claude-code link predates Stage 1; the canonical store both agents link into (~/.agents/skills/) sits outside the backup's scope (~/.claude + ~/.codex only), so the canonical copy's own pre-existence can't be determined either way — kept out of the automated removal rather than guessed at.
  • High — README:57, plugin uninstall can delete persistent data. Added --keep-data to all three claude plugin uninstall lines; confirmed from claude plugin uninstall --help on this host's 2.1.283, and stated what it retains vs. removes. ~/.claude/plugins/data/ (names only) shows context-mode-context-mode/ and codex-openai-codex/ currently exist (so --keep-data preserves real content there); claude-hud-claude-hud/ does not exist today. codex plugin remove --help has no equivalent flag at all — documented that it unconditionally removes both the registration and the cache directory (~/.codex/plugins/cache/context-mode/context-mode/ for the context-mode plugin specifically, confirmed against this same document's own codex -3 receipt excerpt, which retains a command under that exact path).
  • Medium — README:91, workflow removals have no checksum guards. Gated all removable files on a recorded sha256: the 12 workflow files and the 2 existing hooks against the repository's own digest (already established byte-identical, for the hooks; cross-checked SHA256SUMS against the live repo files for the workflows, both match); the 10 agent files and ~/.codex/RTK.md against a digest read from the live file in this pass, explicitly labeled "recorded 2026-09-27 from the live file, not an install-time digest" since no prior byte-identity claim existed for those. contract.config.json pulled out of automatic removal entirely — it's a per-host render with no install-time checksum this document could ever gate against — and left as a documented manual step instead.
  • Medium — Claude receipt -3:48, fixture-read claim overstated. Recorded a superseding mac-coordinator-64gb-20260925--claude-code--use--20260927-4 receipt. Rehearsed by hand twice before recording (once standalone, once as the exact final command text). Its positive control now requires, and prints only a constant sanitized target for (<scratch>/fixture.txt or none, plus booleans — never a real path), a Read tool_use whose own input.file_path, compared by basename only, names the fixture — not merely a Read call present somewhere in the transcript. Added an in-transcript negative control on a wrong basename (notfixture.txt), correctly rejected, alongside the existing wrong-count negative control. Copied -3's other arguments (host, platform, component, stage, evidence class, --second-physical-machine, model, layer-ref); commands 1 and 3 are byte-identical to -3's. The new receipt's recorded_by.identity_sha256 matches -3's and the rest of the chain exactly (91719ef8...), so the README's "every receipt carries the identical identity" claim needed no change. Updated every README reference to the claude-code receipt chain (intro is unaffected; "Host receipts" intro count and table, the --allow-unbound-version sentence, and all three "Limits" bullets) from -3 to -4.

Hot-file protocol: origin/main had moved by 2 commits touching manifests/evidence.json (and, incidentally, 3 of the 10 catalog agent files this PR's README lists digests for: isolated-builder.md, source-scout.md, stack-verifier.md). Did not trust git's clean auto-merge: took main's copy of evidence.json and re-registered this branch's 9 evidence files (8 host receipts + the README) with host_receipts.register_file, then re-ran component_matrix.py --write and new_host_grand_list.py --write (both no-ops — content already matched). Caught and fixed the one place this merge made the README stale: the "matches the repo copy" aside for the agent digests now correctly says 7 of 10 match, not 10 of 10, and explains why (the sha256 gate itself was never affected — it always read the live digest, never the repository's).

Self-review correction (before this comment): re-reading my own README prose against what I'd actually verified turned up three claims that outran their evidence — fixed in a follow-up commit before this one: (1) the skills -a keys were confirmed by reading cli.mjs source, not by triggering a live CLI error as first written; (2) the Codex cache path first named the wrong marketplace (openai-bundled instead of context-mode); (3) the canonical-copy-removal sentence asserted an outcome (detectInstalledAgents() finds nothing else linked) this session never actually ran to confirm — reworded to state the CLI's mechanism instead.

Not fixed, flagged as follow-ups, not in this PR's diff:

  • tools/adoption/install_skills.py's own remove call at line ~196 still uses the unscoped form (no -a); the README now says so explicitly rather than claiming it matches the new scoped rollback.
  • The new -4 receipt's limitation Qualify native recovery and a pinned local GPU workflow #6 says "per this README's ... section" from inside a receipt file, which is a Claude Code session's own README self-reference, not incorrect but a slightly odd phrasing in hindsight — not worth a fifth generation to reword; flagging it here instead.

Checks, all exit 0 at cd0d37eb: host_receipts.py validate (175 receipts, passed), validate.py (7332 hashed files, passed), component_matrix.py --check (32 rows), new_host_grand_list.py --check (32 layers, 66 winners), git diff --check. grep -n "rm -rf\|/Users/" README.md: no hits, both times this was checked (before and after the merge).

🤖 Generated with Claude Code

https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu

@seathatflowsinourveins

Copy link
Copy Markdown
Owner Author

GPT-6 cross-family re-check, round 3, at cd0d37eb: approve. (codex exec -s read-only -m gpt-6-astra, reasoning effort max)

All four round-2 findings are fixed:

  • Skill removal. README.md:275 has exactly 27 scoped removals, the manifest minus typesafe-ai; line 257 documents that skill's preservation and its optional Codex-only removal. Checked against skills v1.7.0 src/remove.ts#L209.
  • Plugin data. --keep-data is on all three Claude uninstalls (README.md:57), and line 156 states what is kept and what is removed. Checked against Claude Code 2.1.283 help and the plugins reference.
  • Checksum guards. All 25 gates (README.md:80, 107, 128, 185) run rm only after equality. 75 synthetic shell cases passed: matching, mismatching and missing hashes per gate. contract.config.json is manual.
  • Fixture-read claim. The claude-code -4 receipt:29 checks the basename of Read.input.file_path and gates its exit on that match, the count, the native exit and both negative controls. Six synthetic checker cases behaved correctly.

Checks: host_receipts.py validate (175) passed; validate.py passed; component_matrix.py --check and new_host_grand_list.py --check exit 0.

Residuals. These are nonblocking and recorded under the bounded-loop rule, with no round 4:

  1. README.md:104 and :123: the digests listed for secret_path_guard.py, test-contract-mutations.mjs and test-envelope.mjs match 327d139d, not HEAD. The guards still protect. Fix later by pinning those source references to 327d139d and noting the later changes.
  2. README.md:244: -a does not have to come last, because the pinned parser stops at the next flag (src/remove.ts#L425), and both orders produced equal options. The listed commands can stay; drop the mandatory-order sentence.

tools/adoption/install_skills.py:196 (the unscoped remove) is correctly a follow-up: the repaired rollback does not call that installer.

🤖 Generated with Claude Code

@seathatflowsinourveins
seathatflowsinourveins merged commit f68d13c into main Sep 27, 2026
24 checks passed
@seathatflowsinourveins
seathatflowsinourveins deleted the claude/mac-stage1-client-layer-20260927 branch September 27, 2026 15:19
@seathatflowsinourveins

seathatflowsinourveins commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner Author

Merged as f68d13c086a381a8bfc7214b90278bf2d03cc3b9 (squash of the approved head cd0d37eb). Before merging, the squash's resulting tree (the head merged onto main a42f0f7f) was validated locally:

  • 0 conflicts;
  • validate.py passed;
  • host_receipts.py validate passed (175 receipts);
  • component_matrix.py --check and new_host_grand_list.py --check exit 0.

Residuals from GPT-6's round-3 review, recorded here as nonblocking. The bounded-loop rule allows no round 4.

  1. README.md:104 and :123: the digests for secret_path_guard.py, test-contract-mutations.mjs and test-envelope.mjs match 327d139d, not the merged head. These files changed on main after the digests were taken. The digest gates still fail closed, so a mismatching file is kept, not removed.
  2. README.md:244 says -a must come last. It need not: the skills parser stops at the next flag. The command form in the README is correct either way.

Follow-up tracked separately: #405, fixed by #407 (install_skills.py's own rollback).

seathatflowsinourveins added a commit that referenced this pull request Sep 27, 2026
…squash

#391's receipts cited commits published only on its head branch; the
repository deletes merged head branches and CI fetches heads and tags
only, so main went red until the revisions were tagged. The 2026-09-25
row now says to tag every such revision (receipt-revision/<sha8>) before
merging and to check it in a full main-plus-tags clone. This PR's own
receipt revision is tagged (receipt-revision/b16b84b8).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
seathatflowsinourveins added a commit that referenced this pull request Sep 27, 2026
manifests/evidence.json and docs/component-evidence-matrix.md taken from
main, this branch's files re-registered, component_matrix and
new_host_grand_list regenerated. validate.py passed (7352 hashed files);
host_receipts.py validate 192 passed; both --check runs exit 0. Receipt
revisions not on main are tagged receipt-revision/<sha8>.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017d27CMzkbLezdLv7vAStPu
seathatflowsinourveins pushed a commit that referenced this pull request Oct 3, 2026
Date the 2026-09-27 Stage 1 narrative (#391) and cite its coverage
section at README.md:414-433 instead of calling it later than the
2026-09-29 check; quote the -76 comment's own words for item 3 instead
of "exploratory"; cite AGENTS.md:12 for supported installation; say the
edition's :101-107 limits, not dates, the agentmemory evidence; scope
the jCodeMunch gap to the Mac.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Oct 3, 2026
…ons (#662)

* docs: preserve PR 508 token-stack history for retirement

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* chore: register PR 508 retirement record evidence

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* docs: apply the PR 662 review round to the PR 508 retirement record

Date the 2026-09-27 Stage 1 narrative (#391) and cite its coverage
section at README.md:414-433 instead of calling it later than the
2026-09-29 check; quote the -76 comment's own words for item 3 instead
of "exploratory"; cite AGENTS.md:12 for supported installation; say the
edition's :101-107 limits, not dates, the agentmemory evidence; scope
the jCodeMunch gap to the Mac.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore: re-register the PR 508 retirement record after the review round

Take main's manifests/evidence.json and re-register only the edited
record (sha256 7227ae70...73a7, 21,068 bytes) under the hot-file
protocol; both report generators ran and changed nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: call the PR 508 merge trigger dormant, not void, in the retirement record

Review thread on #662 (record line 220): the #508-merges branch of the
overturn condition at docs/decisions/2026-09-30-task-model-routing.md
:213-216 cannot fire while #508 stays closed, but the record keeps #508's
branch and head, and a closed pull request can be reopened (GitHub GraphQL
reopenPullRequest, docs.github.com/en/graphql/reference/pulls
#mutation-reopenpullrequest; live schema: "Reopen a pull request."), so a
reopened #508 that merges would fire it. Scope ast-grep's :217-218
admission condition to "while #508 stays closed", since routing :201
also ties a Gate A path to #508 merging as written.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore: re-register the PR 508 retirement record after the dormant-trigger fix

Re-register only the edited record on the branch's own
manifests/evidence.json (sha256 59f58e14...c70d, 21,398 bytes); main is not
merged. scripts/component_matrix.py --write and
scripts/new_host_grand_list.py --write ran and changed nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Scout <scout@local>
Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant