Skip to content

Fix native worker timeout cleanup on macOS - #1

Merged
seathatflowsinourveins merged 1 commit into
mainfrom
codex/mac-worker-timeout-current
Sep 20, 2026
Merged

seathatflowsinourveins merged 1 commit into
mainfrom
codex/mac-worker-timeout-current

Conversation

@seathatflowsinourveins

Copy link
Copy Markdown
Owner

A timed-out native research worker retired its process group twice. On macOS, the second signal could raise PermissionError and hide the timeout result. This change performs cleanup once in the shared finalization path and preserves the worker's exit status.

The descendant checks now use POSIX process state so they actually verify cleanup on both macOS and Linux.

Validation: reproduced the Mac failure before the fix; all 13 targeted worker tests pass, including cancellation and a descendant that ignores TERM. The current full Mac suite passes 259 tests with 34 optional SDK tests skipped, using a canonical temporary directory outside Git. Both repository validators and the redacted secret scan pass.

Retire timed-out process groups once in the shared cleanup path so a second signal cannot mask the timeout result on macOS. Preserve the child exit status, verify timeout and normal-exit behavior, and use POSIX ps for descendant checks on macOS and Linux.

Update the two covered source hashes in the evidence manifest.
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-20T02:05:35.766279Z ce39b3c PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@seathatflowsinourveins
seathatflowsinourveins merged commit 21e6836 into main Sep 20, 2026
4 checks passed
seathatflowsinourveins added a commit that referenced this pull request Sep 23, 2026
…okenizer, exact hostname) (#104)

* Fix CodeQL first-analysis alerts: href scheme allowlist, case-insensitive tag scan, exact hostname check

Resolves the 5 real defects from the repository's first CodeQL default-setup
analysis (commit 168a3a8, alerts 1/3/4/5/8), re-located at base 796f759 since
PR #96 changed the generated explorer between the two:

- js/xss-through-dom (docs/ecosystem/template.html:145): link() now builds
  href through a safeHref() helper that only allows http:/https: URLs,
  blocking a javascript:-URI href from catalog data.
- py/bad-tag-filter (scripts/build_ecosystem.py:704,
  tests/test_ecosystem_manifest.py:231, tests/test_claude_repository_evidence.py:147):
  add re.IGNORECASE so an injected uppercase <SCRIPT> tag is still counted
  into the page's inline-script CSP hash instead of silently bypassing the
  single-script precondition.
- py/incomplete-url-substring-sanitization (tests/test_lifecycle_capture.py:103):
  replace the "sec.gov" in url substring check with an exact
  urlparse(url).hostname comparison.

The remaining 5 alerts (2, 6, 7, 9, 10) are false positives / test-only
synthetic-secret fixtures; their justification is recorded for the
coordinator to apply via the GitHub API in
docs/decisions/2026-09-22-codeql-first-analysis.md and the sibling
codeql-dismissals.json, not dismissed here.

manifests/evidence.json's sha256/bytes entries for the five touched files
are refreshed so scripts/validate.py (run by validate.yml) still passes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Correct alert 7/9 location prose and add an executed safeHref check (review fixes)

Resolves the independent reviewer's three medium findings on 65f73e2:
alerts #7 and #9's location descriptions and dismissal comments pointed at
the wrong code after the 796f759 re-location (both now cite the actual
CodeQL sink); alert #1's "node -e smoke check" claim named no runnable
command, so it is replaced with an executed unittest that runs the
committed safeHref helper (extracted verbatim from template.html) under
Node and asserts javascript:/data:/vbscript:/file:/mailto: are rejected
while http(s) and relative URLs pass through, and alert #1 is relabeled
defense-in-depth given build_ecosystem.py's public_url() is the primary
control. Refreshes manifests/evidence.json for the two touched files so
validate.py stays green. Re-running the full suite three times confirms
the reviewer's flagged skip-count (338) is deterministic, not a flake.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Replace the <script> regexes with stdlib HTML tokenizers (CodeQL py/bad-tag-filter)

re.I alone would likely leave py/bad-tag-filter open (it also flags missed
end-tag variants such as </script >). The build script and both tests now use
html.parser, which tokenizes like a browser; on the real generated pages the
parsers return the identical single body the regexes returned. Corrects the
record's alert #1 control description (loopback_url also admits http loopback
links) and the skip-count claim.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fail loudly where html.parser and a browser disagree on <script>

Review follow-up: a self-closing <script/> or a script after <!--> (Python 3.12)
was invisible to InlineScripts. render_from_data now requires no self-closing
script and requires the parser's script-start count to equal the raw
"<script" count; the CSP test asserts the same count. The record's loopback_url
and browser-equivalence wording is corrected.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: seathatflowsinourveins <234074349+seathatflowsinourveins@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 23, 2026
…ng for bootstrap

Codex's round-3e review of c18aea3 found two new Highs and four Mediums --
each earlier round's hand-written rollback had closed one edge case and
opened another (an unowned rm -rf, a partial hard-copy backup, a reload
skipped at a step boundary, cleanup left armed after success, a signal
racing a foreground mv, an unguarded rm aborting recovery under set -e).
Per the coordinator's explicit direction, this round changes the approach
instead of patching further: simplify toward idempotent convergence, the
way `brew services` works, with step tables and signal-injection tests
proving every row.

Both scripts now explicitly trap INT, TERM and HUP (`exit 130/143/129`
only; recovery logic stays in the EXIT handler alone) -- bash only
reliably defers a caught signal to a completed step when the signal is
explicitly trapped, not left at its default disposition (confirmed
directly: an untrapped SIGINT sent to a bash script with only an EXIT trap
registered does not interrupt it at all in this environment, while an
untrapped SIGTERM does -- the asymmetry Codex's finding 5 was pointing at).

bootstrap-macos.sh (High #1, unowned rm -rf): new prune_old_version()
deletes a superseded versioned directory only when its canonical form is a
direct child of the canonical tools/ directory with a name matching
<id>-<version>-<stamp>; an external target, a relative "../" target, a
non-matching name, or a symlink loop (canonical_path's python3
os.path.realpath handles cycles without hanging) is logged and left in
place, never deleted. Pruning is now genuinely best-effort: a rejected or
failed deletion warns and never fails the install. Four new tests, one per
adversarial target shape, confirmed to fail against c18aea3 (the function
does not exist there).

launchd-agents.sh -- full redesign, net roughly flat in size (+22 lines)
despite ~390 changed, since one pure function replaces four tracking
markers (pending_backup_plist/_dest/_reload_needed, pending_dest_tmp) and
their two independent, historically drifting recovery implementations:
- backup_plist is a hard link (`ln`, High #2), not `cp`: it either exists
  completely or not at all, so a partial copy can never overwrite the
  intact original. Same directory as dest_plist (round 3d), never this
  script's own state_dir.
- Staging (cp to dest_plist.new, then mv over dest_plist) is unchanged in
  shape but the name lost its PID suffix, matching the backup's own
  naming, so a stale artifact from a crashed run is visible to and
  reconciled by the very next run, not just the same process.
- reconcile_install replaces every marker: it reads dest_plist's identity
  against backup_plist (-ef, since the backup is a hard link) and
  launchd's own current loaded state, and converges to whichever ONE
  action that state calls for (drop the backup, reload, restore-and-
  reload, or -- if nothing converges -- keep the backup and name the exact
  re-run command). It is the ONLY recovery logic: called explicitly after
  a normal attempt, and identically as the EXIT trap for anything that
  interrupts one. `set +e` inside it, explicit checks throughout, per the
  coordinator's instruction that the handler must never itself become
  fatal under the script's own `set -e`.
- record_enabled_label now runs as soon as install actually commits to a
  label (not only on success), so a label that ends up needing attention
  is never re-refused as unowned on retry; a was_already_enabled snapshot,
  taken before that point, keeps the L2 bootout gate correctly checking
  ownership as it stood BEFORE this run, not after.

New tests: a shared _signal_shim helper plus a parameterized step x signal
matrix (ln and mv, TERM and INT) proving retry-converges for launchd; a
dedicated SIGTERM/SIGINT-parameterized migration test and three more for
bootstrap-macos.sh's prune_old_version. Every new/changed test confirmed
against c18aea3 by stashing the fix; several rewritten round-3b/c/d tests
(now asserting the new design's own honest behavior -- e.g. a fresh
install that never loads is left on disk with a manual-retry message,
never silently deleted) checked to still exercise the coordinator's
original findings under the new mechanism.

Full acceptance re-run on this commit: full suite (2821 tests) passes
except the one evidence-registration lag scripts/validate.py itself always
flags immediately after an edit (resolved by the registration commit that
follows); named macOS/launchd modules (126 tests) and the bash-3.2/
launchd-gated subset (56 tests) under a real compiled GNU bash 3.2.0 all
pass; --plan clean under both bash versions; shellcheck and zizmor clean;
git diff --check clean; guarded gitleaks clean. origin/main had not moved
since it was already merged into c18aea3's own parent, so no new merge was
needed this round. No macOS host ran any of this.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 23, 2026
…hd agents (brew-services semantics), embedding acceptance, recording-tooling CI (#94)

* Pin socraticode and platform-dependency integrity, add launchd agents for the macOS adoption profile

adoption/bootstrap-macos.sh: replace the jq-only Homebrew block with a
bash-3.2-safe loop over the six floating formulae (jq, python@3.13, ripgrep,
coreutils, restic, shellcheck), installing only what `brew list --versions`
reports missing, gated on --skip-system-packages/--plan, and echoed in
--plan. socraticode now has a reviewed npm pin (--ignore-scripts, matching
the Linux recipe) instead of being a documented, unpinned skip, so
documented_unpinned_ids is now empty (guarded for bash 3.2). Add
verify_platform_dependency: a fail-closed check, run right after `npm
install`, that the darwin-arm64 platform_dependency npm resolves for codex
and claude-code (name/version/integrity) matches the host's own
.package-lock.json or package.json _integrity.

adoption/pins-macos-arm64.json: add the socraticode pin (sha256 verified by
downloading the tarball and cross-checking against `npm view
socraticode@1.14.0 dist.integrity`) and a platform_dependency entry for each
of @openai/codex and @anthropic-ai/claude-code, with npm dist.integrity read
2026-09-23. @openai/codex-darwin-arm64 is not itself a published package
name; `npm view @openai/codex@0.155.1 optionalDependencies` shows it as an
npm: alias to @openai/codex@0.155.1-darwin-arm64, recorded as such.

New adoption/launchd/: three plist templates (qdrant, ai-memory, the
llama.cpp Metal embedding server on port 8232) using the same ${NAME}
placeholder convention as adoption/templates/, plus launchd-agents.sh
(render/lint/install/status/remove; plutil -lint with a plistlib fallback;
launchctl bootstrap/print/bootout; remove only boots out a label this
script's own state file recorded as enabled, never deleting data). New
tools/adoption/render_launchd.py reuses render_config.py's render_one/
load_host_values/parse_set_values. New adoption/hosts/macos-example.json
supplies the render_config placeholder schema with /Users/example paths and
cites hardware-profiles.json's macos-arm64-48gb-projected id; EMBED_URL is
the bare host:port form (not the http://-prefixed form first drafted) since
project.codex.config.template.toml already prepends http:// for it.

.github/workflows/adoption-bootstrap.yml: add tools/adoption/** to the path
filters; add a launchd render/lint/bootstrap-qdrant/health-check/bootout
step to the existing bootstrap-macos job; add bootstrap-macos-brew (runs the
script without --skip-system-packages, asserting python3.13/rg/restic/
shellcheck/gsha256sum) and validate-macos (the repository's own validators,
the adoption test modules as a gate, and one full python3 -m unittest
recorded as a non-gating artifact), both on macos-15.

adoption/platforms/macos-arm64.md: prerequisites, exit-code table,
darwin-arm64 pin table and launchd section updated to match; the
"unpinned"/documented-skip wording for socraticode is removed since it is
now pinned.

tests/test_adoption_bootstrap_macos.py: update DOCUMENTED_SKIPS, the brew
block assertions, and add PlatformDependencyVerificationTests (offline,
fixture pins + fake .package-lock.json/npm shim; fixture integrity values
are assembled from concatenated parts, not written as single literals).
New tests/test_adoption_launchd.py covers template rendering/schema,
render_launchd.py, launchd-agents.sh structure and an offline
render/lint/install/status/remove cycle against a fake launchctl.

Evidence class: no macOS host ran anything here. Everything is either local
integration (Linux-side unit tests, shellcheck, plistlib parsing) or a
hosted macos-15 runner the coordinator dispatches later.
adoption/platforms/macos-arm64.md stays drafted_not_accepted, and
adoption/manifest.json is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Re-hash the evidence registry and rebuild the ecosystem explorer

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fix round: real npm/lockfile evidence, bash 3.2 consumption bug, launchd ownership, XML escaping, harden-runner

Addresses the Opus evidence review (fail), the Codex cross-family review
(fail), and hosted macos-15 CI run 35812470345 on PR #94, which reproduced
three of these findings directly and one more (git_revision) on its own.

1. npm platform-dependency verification never could have passed: measured
   directly against npm 11.19.0 on this project's own host (existing
   ecosystem prefixes, e.g. mcporter-0.13.13, and a fresh `npm install
   --global --prefix <dir> is-odd`), a --global --prefix install writes no
   lockfile anywhere under the prefix, no package.json _integrity, and
   places a platform-specific optional dependency NESTED under the parent
   (lib/node_modules/<parent>/node_modules/<platform-dep>), not at its own
   top-level name -- reproduced here with two real npm-packed local
   packages (tests.RealNpmLockfileEvidenceTests), not asserted from memory.
   verify_platform_dependency (post-hoc lockfile/package.json read) is
   replaced by install_platform_dependency: fetch the platform tarball
   itself (dist.tarball, added to each platform_dependency pin alongside a
   freshly re-hashed sha256), verify it with the same fetch() every other
   pin uses, then `npm install --global --prefix <dir> <alias>@file:<path>`
   the already-verified local archive -- confirmed with real npm
   (PlatformDependencyInstallTests) that this places the verified bytes at
   the alias name regardless of the tarball's own internal package.json
   name, which is what Node's require() actually resolves by.

2. bootstrap-macos.sh:192's `for allowed_id in "${allowed_unpinned_ids[@]}"`
   was unguarded: the array's assignment was bash-3.2-safe but this later
   consumption of it was not, and it aborts with "allowed_unpinned_ids[@]:
   unbound variable" whenever it is genuinely empty -- exactly the hosted
   CI failure. Fixed with the same ${arr[@]+"${arr[@]}"} guard used
   elsewhere. This fix round compiled a real bash 3.2.0 from source
   (GNU bash source, ftp.gnu.org) to reproduce the abort and confirm the
   fix dynamically -- not just structurally, which is all bash 5's relaxed
   nounset handling of empty arrays can ever show. ScriptBehaviorUnderRealBash32Tests
   runs the exact --plan and exit-3 paths under that real 3.2 binary when
   BASH32_BINARY is set, or on a real Mac (/bin/bash genuinely is 3.2 there).

2b. tests/test_adoption_launchd.py's end-to-end cycle test asserted the
   plistlib fallback text unconditionally, which fails wherever plutil is
   present (a real Mac; not this dev host). Relaxed to a generic success
   check, with two new dedicated, PATH-controlled tests exercising each
   path deliberately (test_lint_falls_back_to_plistlib_when_plutil_is_absent,
   test_lint_prefers_plutil_when_present), so the behavior is proven either
   way rather than incidentally on whichever host happens to run it.

2c. scripts/adoption_status.py's git_revision(root) compared git's own
   resolved --show-toplevel path against an unresolved caller root; macOS
   routes its default tempdir through /var -> /private/var (reproduced here
   with a plain symlink on Linux), so a direct caller with an unresolved
   root -- exactly what the existing unit test does -- got None instead of
   the real SHA on a real Mac. inspect_adoption already resolves root
   first, so this was latent; git_revision now resolves its own root
   parameter, closing the gap for every caller.

3. launchd-agents.sh install now refuses to overwrite an existing
   ~/Library/LaunchAgents/<label>.plist that this script did not itself
   record as enabled, and creates state/logs and state/qdrant (qdrant's new
   WorkingDirectory) before installing -- launchd does not create missing
   parent directories for StandardOutPath/StandardErrorPath. remove keeps
   ownership (does not forget the label) when launchctl bootout fails, and
   on success now deletes the copied plist under ~/Library/LaunchAgents
   (not just booting it out) so RunAtLoad cannot reload it at the next
   login; it still never touches the component's own data or logs.
   llama-embed is left out of the default install/status/remove set until a
   model argument exists (render/lint still cover it, and an explicit
   --label still installs it).

4. tools/adoption/render_launchd.py rendered plists via a raw
   string.Template substitution on template text, which is not
   XML-escaping: a substituted value containing "&" (a real path such as
   /Volumes/R&D/eco) would corrupt the document. Rendering now parses the
   template as a plist first (plistlib tolerates the still-literal ${NAME}
   placeholders as ordinary string content), substitutes inside every
   string leaf of the resulting structure, and re-serializes with
   plistlib.dumps, so plistlib's own escaping covers every value; proven
   with a real "&"-containing host value round-tripped through valid XML.

5. The bootstrap-macos-brew and validate-macos jobs, and the existing
   bootstrap-macos job (which had none), now start with harden-runner in
   audit mode: the pinned version's own README (read at its exact pinned
   commit) states macOS runners get full audit-mode support, which is the
   only mode this project ever uses. tests/test_workflow_hardening.py's
   blanket "macOS is exempt" rule is narrowed to a named,
   ownership-scoped exemption for hardware-profile-smoke.yml only (owned by
   a different, concurrent task and off limits to edit here), with a new
   regression test enumerating this branch's three macOS jobs explicitly.

6. bootstrap-macos-brew no longer runs actions/setup-python (which could
   shadow python@3.13 with a GitHub-provided interpreter on PATH, making
   the check pass without ever exercising Homebrew's own install); it now
   checks a short, explicit list of brew --prefix candidates for
   python3.13 and fails with a clear message (not a false pass) if none
   resolve.

7. Path filters gained tests/test_adoption_*.py and the five scripts
   validate-macos runs.

8. The launchd smoke step's "qdrant binary absent" branch exited 0; qdrant
   is a pinned, always-selected component (not itself in required_commands,
   which is owned by adoption/manifest.json and not edited here), so this
   now exits 1 -- its absence after a successful bootstrap is a real
   failure to surface.

9. adoption/platforms/macos-arm64.md's two previously-cited green hosted
   runs are now marked as predating this fix round, naming the actual
   failing run (35812470345) and what it caught, so a reader does not treat
   stale hosted evidence as covering the current script.

10. brew_formulae's install order is now jq alone (needed by the
   --profile/pins validation that follows), then that validation, then the
   remaining five formulae -- previously all six installed before any
   validation ran.

Evidence class: still no macOS host ran anything here. The npm/lockfile
finding and the bash 3.2 abort are now independently reproduced (real npm,
a real compiled bash 3.2.0, and a real /var-style symlink) rather than
argued from memory; the launchd ownership/XML-escaping/harden-runner fixes
remain offline local_integration and structural evidence pending the next
hosted macos-15 run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Rehash the evidence registry via host_receipts.register_file and rebuild the explorer

Used scripts/host_receipts.py's register_file (an absolute Path.cwd() root,
per the coordinator's fix-round instructions) for every file this fix round
changed, instead of hand-patching manifests/evidence.json: the six already
registered in the previous round's rehash (adoption-bootstrap.yml,
bootstrap-macos.sh, pins-macos-arm64.json, macos-arm64.md,
test_adoption_bootstrap_macos.py) plus scripts/adoption_status.py, and this
round's four new files that were not previously registered
(com.native-stack.qdrant.plist.template, launchd-agents.sh,
test_adoption_launchd.py, test_workflow_hardening.py, render_launchd.py).
register_file rewrites the whole manifest via json.dumps(indent=2), so this
diff also reformats entries this round did not otherwise touch; the content
change is exactly the updated/added hash and byte-count records.

docs/ecosystem/index.html: regenerated with `scripts/build_ecosystem.py
--write` from the manifests above and registered the same way; --check
passes.

python3 scripts/validate.py, scripts/validate_catalogs.py,
scripts/validate_foundation.py --root . --json, scripts/landscape.py --root
., scripts/build_ecosystem.py --check and
tools/sota-convergence/build_verdicts.py --check all pass after this commit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Round 2: defeat the npm shadow-copy, launchd orphan/unloaded/directory fixes, remaining lows

Addresses the Opus round-2 re-review (fail): one High, one Medium, several
Lows. Full evidence for each below.

**High: npm's own unverified fetch shadowed the verified platform tarball.**
install_npm installed the wrapper without deferring its own optional-dependency
resolution, so npm auto-fetched the darwin-arm64 platform package itself,
unverified, and nested it under the wrapper's own node_modules -- and Node's
require.resolve checks that nested copy before a top-level sibling, so the
previous round's `<alias>@file:<path>` top-level install was never what
codex's bin/codex.js or claude-code's install.cjs actually resolved.
Measured directly (npm 11.19.0) that neither --omit=optional, --no-optional
nor NPM_CONFIG_OMIT=optional suppress the physical on-disk fetch for this
install shape (a single global tarball install, no lockfile), and that
pre-placing verified content at the nested path before installing the
wrapper does not survive it either (npm overwrote it, "changed N packages").
install_platform_dependency now: installs the wrapper with --ignore-scripts
first (deferring its lifecycle scripts); asks Node itself, via
require.resolve scoped to the wrapper's own directory, exactly where it
would resolve the dependency from; deletes that path and extracts the
independently SHA-256-verified tarball there (falling back to a top-level
alias only if Node found nothing to shadow); re-verifies resolution and the
resolved version against the pin, fail closed on either mismatch; then
install_npm runs `npm rebuild --global` for the wrapper, npm's own
documented way to run the lifecycle scripts an --ignore-scripts install
deferred, now against the verified copy. Proven end to end with real npm
against two fixtures mimicking the two real shapes: a codex-like wrapper
(no lifecycle scripts, resolves at every invocation) and a claude-code-like
one (a postinstall that copies from wherever require.resolve finds the
platform package into its own bin/ file, exactly like install.cjs) --
tests.PlatformDependencyInstallTests.test_codex_like_wrapper_resolves_the_verified_dependency_not_the_shadowed_one
and .test_claude_code_like_wrapper_postinstall_copies_the_verified_dependency,
each asserting the FINAL resolved/copied content is the verified fixture
payload, never the differently-content "unverified" one an attacker-controlled
registry fetch stands in for.

**Medium: launchd-agents.sh install left an orphaned, unowned plist on a
failed bootstrap.** cp ran before launchctl bootstrap; a failed bootstrap
left the copied plist under ~/Library/LaunchAgents where RunAtLoad would
reload it at the next login, and neither a retry (the ownership check) nor
remove (never touches an unrecorded label) could ever reach it again. Now
rolls the copy back (rm -f the destination) and exits 1 on a failed
bootstrap, never recording it as enabled.

**Lows:**
- remove on an owned-but-unloaded label always failed calling bootout on
  something not loaded; now checks `launchctl print` first and, if not
  loaded, cleans up (deletes the plist, forgets the label) directly instead.
- install's directory creation was hardcoded to this shell's own
  ECO_INSTALL_ROOT (state/logs, state/qdrant), which breaks a plist rendered
  with --host against a different host's ECO_ROOT value. Now reads each
  rendered plist's own StandardOutPath/StandardErrorPath/WorkingDirectory
  (via a new plist_value helper, plutil-or-plistlib-fallback like cmd_lint)
  and creates directories from those declared paths instead.
- tests/test_workflow_hardening.py's MACOS_JOBS_OWNED_ELSEWHERE exempted the
  whole hardware-profile-smoke.yml file; keyed as
  "hardware-profile-smoke.yml:macos-profile" now, so that file's own ubuntu
  linux-profile job stays checked (new regression test confirms it does).
- Path filters gained scripts/adoption_status.py in both push/pull_request
  lists.
- render_launchd.py's render_plist now also catches ValueError from
  string.Template.substitute (a malformed placeholder, e.g. a bare trailing
  "$") and raises it as the same RenderError every other failure uses,
  instead of an uncaught traceback.
- scripts/adoption_status.py's git_revision now also catches RuntimeError,
  which Path.resolve() raises on Python before 3.13 for a symlink loop
  (3.13+ instead returns the unresolved remainder); reproduced with a real
  two-node symlink loop on this host's Python 3.12.

Also fixed in passing: tests/test_adoption_bootstrap.py's
test_workflow_installs_the_manifest_supported_python_before_status did a
whole-file text.index("scripts/adoption_status.py") search, which this
round's new path-filter entry (a bare, unquoted-command list item earlier in
the file than any job's real invocation) made match before the intended
"actions/setup-python@" occurrence; narrowed the search to the literal
"python3 scripts/adoption_status.py" invocation.

Evidence class: the npm shadow-copy fix and the symlink-loop fix are real
npm/Python behavior, measured and reproduced directly on this host (no
macOS host ran any of this); the launchd fixes remain offline
local_integration and structural evidence pending the next hosted macos-15
run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Round 3: reinstall-while-loaded, arbitrary-path deletion, npmrc/postinstall gaps, atomic staging

Addresses the Opus pass-with-findings review of round 2 (1c01f4f): the High
(npm shadow-copy) is confirmed closed by that review (real tarballs
downloaded, all four hashes confirmed, tarball contents and npm rebuild
scope checked, prepare confirmed not to run). This commit fixes the
remaining Medium/Low/Nit findings.

1. Medium, launchd-agents.sh install: re-running install on a label that is
   owned AND already loaded (e.g. after a re-render) had cp overwrite the
   live plist in place, then launchctl bootstrap fail (already bootstrapped),
   then the rollback delete the just-overwritten file out from under the
   still-running service. install now checks launchctl print first and
   boots the label out before reinstalling if it is loaded (refusing and
   leaving it running if that bootout itself fails); it only ever deletes
   the destination on a failed bootstrap when it did not already exist,
   otherwise it backs it up first and restores (and best-effort reloads) it.
   New shim test with a stateful launchctl (bootstrap sets a per-label
   loaded marker, print reads it, bootout clears it) drives install into
   exactly this already-loaded-reinstall-then-bootstrap-fails path and
   checks the original plist survives byte for byte.

2. Medium, bootstrap-macos.sh install_platform_dependency: `rm -rf` ran on
   whatever path require.resolve returned. Measured directly: Node's
   require.resolve, even given an explicit `paths` array, still searches its
   GLOBAL_FOLDERS fallback (NODE_PATH entries, $HOME/.node_modules, etc. --
   Node's own module docs); setting NODE_PATH to a decoy directory made it
   resolve OUTSIDE the prefix entirely, and the previous code would have
   deleted that decoy -- a real, exploitable arbitrary-path deletion, not a
   theoretical one. The resolved path is now accepted only when it equals
   the exact expected nested path
   ($prefix/lib/node_modules/$package/node_modules/$dep_name/package.json);
   anything else falls back to the existing top-level alias instead of
   touching whatever require.resolve actually named. After installing the
   verified copy, the script now also asserts the resolved path equals
   $target_dir/package.json and that the installed package's own `name`
   field equals the pin's `resolved_package` (pinned since round 1, unread
   by any check until now), fail closed on either mismatch. New tests: a
   real decoy directory via NODE_PATH that survives completely untouched
   while the verified content still lands via the top-level fallback, and a
   resolved_package mismatch that fails closed even though the sha256 and
   version both verify.

3. Low: npm rebuild now passes --ignore-scripts=false explicitly, since a
   user-level .npmrc with ignore-scripts=true would otherwise make it return
   early without running anything. Added `postinstall_binary_check` (an
   optional pin field: platform_file / wrapper_file) and a cmp -s comparison
   after npm rebuild for claude-code specifically (its install.cjs copies
   package/claude, confirmed by downloading and `tar -tzf`ing the real
   2.1.278 darwin-arm64 tarball, into bin/claude.exe): the check fails
   closed if what actually landed in bin/ does not byte-for-byte match the
   already-verified platform dependency, rather than trusting npm rebuild's
   own exit code alone. New tests cover both the matching case and a
   fabricated mismatch (a postinstall that writes fixed wrong content
   instead of resolving the platform dependency).

4. Low: a re-run of install_npm for the same id/version into an
   already-populated final_prefix would install directly into the live
   prefix, leaving a window where the wrapper's freshly updated files point
   at whatever platform dependency npm's own install just auto-fetched,
   unverified, before the fix-up runs. install_npm now stages into a fresh
   directory under stage_dir when a platform_dependency is involved and
   moves the finished, fully verified result into the final prefix only
   after install_platform_dependency succeeds (mv on the same filesystem is
   a single atomic rename), so the live prefix is always either the
   complete old install or the complete new one. New test runs install_npm
   twice for the same id/version and checks both succeed, the staging
   directory never survives either run, and the final content is still
   fully verified after the re-run.

5. Low, launchd-agents.sh remove: a launchctl print failure was treated as
   "not loaded" for ANY nonzero exit; now only 113 ("Could not find
   service") is treated that way. Any other print failure (permission,
   launchd itself unresponsive, ...) keeps ownership and reports, rather
   than guessing either way. The shared _launchctl_shim test helper was
   upgraded to track loaded state per label (bootstrap sets a marker, print
   reads it, bootout clears it) to match this more realistic behavior, and
   every affected test's expected launchctl-print/bootout counts were
   recomputed and verified against the actual, now-more-realistic call
   sequence.

6. Low: pins-macos-arm64.json's top-level `source` field and
   tests/test_adoption_bootstrap_macos.py's
   test_client_npm_pins_name_their_unpinned_darwin_arm64_dependency still
   described/named the round-1 `<alias>@file:<path>` design and the
   "unpinned" framing. Both updated; the test (renamed) now also asserts
   each platform_dependency carries a real sha256, not just a mention in
   install_note.

7. Nit: render_launchd.py's plistlib.dumps call moved inside the same try
   as _substitute, so a control character in a substituted value (which XML
   plist text content cannot represent) raises the same RenderError every
   other failure here does, instead of an uncaught ValueError. New test
   confirms this with a real control character.

Evidence class: the shadow-copy defeat (closed by the coordinator's own
independent Opus re-review, which downloaded and hashed the real tarballs
itself), the NODE_PATH decoy, the resolved_package check and the atomic
staging are all measured against real npm/Node on this host; the launchd
ownership/reinstall fixes remain offline local_integration and structural
evidence pending the next hosted macos-15 run (hosted CI at 522bbe3 is
already green, 15/15, per the coordinator).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Regenerate the new-host grand list after #101/#102 and re-register round-3 evidence

Reflects the macOS pin/socraticode changes in catalogs/landscape/new-host-grand-list.json
and docs/new-host-grand-list.md (scripts/component_matrix.py --write produced no diff;
scripts/new_host_grand_list.py --write did). Also re-registers the six files changed by
round 3 (bootstrap-macos.sh, launchd-agents.sh, pins-macos-arm64.json,
test_adoption_bootstrap_macos.py, test_adoption_launchd.py, render_launchd.py), which had
been left registered at their pre-round-3 hashes after the origin/main merge. Verified via
scripts/component_matrix.py --check, scripts/new_host_grand_list.py --check,
scripts/evidence_manifest.py --check and scripts/validate.py, all passing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Round 3b: canonicalize paths across a symlinked ancestor; launchd install ordering/gating

bootstrap-macos.sh (Opus M1 / hosted macos-15 run 35820422561; Codex round-3
non-atomic-swap Medium):
- Add canonical_path(), a bash 3.2-safe, realpath(1)-free canonicalizer
  (python3 os.path.realpath, falling back to cd -P/pwd -P and a manual
  ancestor walk for a path that does not exist yet). A real Mac's /var is a
  symlink to /private/var, and Node's require.resolve realpath-resolves
  symlinks by default when it locates a module, so a prefix built under a
  symlinked ancestor (TMPDIR-based staging, or a test fixture root) never
  string-equaled the hand-built expected path. install_platform_dependency
  now canonicalizes its own prefix argument on entry and the resolved path
  Node reports; the post-install resolution assertion also realpaths both
  sides via fs.realpathSync. install_npm canonicalizes final_prefix and the
  staged/live prefix. Reproduced off-Mac with two new tests that build the
  fixture prefix/ecosystem_root under a symlinked tmp root (mirroring /var ->
  /private/var); both fail with "function canonical_path not found" (or, pre
  this fix, the exact "resolved to /priva..." mismatch) without the fix.
- L1: replace the non-atomic `rm -rf "$final_prefix"; mv` prefix swap with
  rename-old-aside, move-staged-in, then delete-old -- a crash between the
  two never leaves final_prefix's own bin_dir symlinks pointing at nothing
  with the previous, fully verified install unrecoverably gone.
- N2: rewrite install_platform_dependency's own design comment to state the
  exact-match containment condition and the resolved-package-name check, and
  the new canonicalization; update macos-arm64.md's platform_dependency
  paragraph (Codex), replacing its stale description of a lockfile/
  `_integrity`-based verifier that no longer exists.

launchd-agents.sh (Opus + Codex round-3):
- L2: gate the pre-reinstall launchctl print/bootout probe on
  is_enabled_label, so install never unloads a label it does not own merely
  because it happens to be loaded.
- L3: move directory creation and the owned-destination backup before that
  probe, so a failure in either never leaves an already-unloaded service
  down with nothing done to restore it.
- L4: add wait_until_unloaded, a bounded poll of launchctl print (checked
  before any sleep, so the common case costs no wall-clock time) after a
  successful bootout, since bootout can return before real launchd teardown
  finishes; refuse rather than bootstrap over a service that never actually
  reported unloaded.
- L5: install now treats only launchctl print exit 113 as confidently "not
  loaded", matching remove; any other failure refuses instead of silently
  assuming unloaded.
- L6: add explicit remove coverage for print exiting 5 (keeps ownership and
  the plist; the existing code path was already correct, now asserted).
- N1: move the temporary pre-overwrite backup out of
  ~/Library/LaunchAgents (which launchd auto-loads at login) into
  state_dir/backups, with a trap-based cleanup (pending_backup_plist) as the
  safety net for a premature exit.
- Five new behavior tests plus three structural ones; three of the five
  (L2, L4, L5) fail against the pre-fix script, confirmed by stashing it.

Full acceptance re-run on this commit: named macOS/launchd modules and the
full suite (2688 tests) under bash 5, the same bash-3.2/launchd-gated subset
(50 tests) under a real compiled GNU bash 3.2.0, --plan clean under both bash
versions (shimmed Darwin/arm64), shellcheck and zizmor clean, git diff
--check clean, scripts/validate.py and the catalog/foundation/component-
matrix/grand-list/evidence-manifest checks all passing, and guarded gitleaks
(dir scan and the git-history scan scoped to this branch's own commits since
origin/main) finding no leaks. No macOS host ran any of this.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Round 3c: rollback the failed prefix swap; stop the launchd trap truncating the live plist

bootstrap-macos.sh (Codex round-3 verification of 5310fe5, Medium 1):
- The rename-old-aside/move-staged-in swap had no rollback if the second
  move (staged prefix into final_prefix) itself failed: final_prefix was
  left absent with every bin_dir symlink into it broken, even though the
  previous, fully verified install still existed on disk as
  final_prefix.previous.$$. Failure injection (a shimmed mv failing on any
  "*-staged" source) confirmed this on 87327dc: exit 1, final_prefix
  absent, only the ".previous." copy remaining. Fixed by moving that
  previous copy straight back into final_prefix's place before returning,
  so a failed move loses only the newly staged prefix, never the working
  one bin_dir already pointed at.
- Corrected the design comment's "atomic" claim (also flagged) to say what
  is actually guaranteed: this swap is recoverable, not atomic; a
  genuinely atomic swap would need a versioned directory plus a single
  symlink flip, a larger restructuring this fix does not make.
- New failure-injection test proves the live prefix (and require.resolve's
  own resolution into it) survives a failed final move with the original
  content intact; confirmed failing against the pre-fix script by stashing
  it (reproduces the exact "absent, previous-only" outcome Codex reported).

launchd-agents.sh (Codex round-3 verification of 5310fe5, Medium 2):
- The round-3b N1 trap was itself a regression: `cp -- "$source_plist"
  "$dest_plist"` wrote directly into the live destination, so a copy that
  failed partway (disk full, an interrupted write) could truncate it and
  exit via `set -e` before any of cmd_install's own restore branches ran --
  and the trap only DELETED the pending backup instead of restoring it,
  losing both the live plist's content and its one recovery copy (the
  pre-N1 script at least kept the backup in this scenario). Differential
  injection (a shimmed cp truncating any destination under
  .../LaunchAgents/*.plist) confirmed the regression against 87327dc's
  pre-3c script: destination truncated, backup deleted.
- Fixed two ways: (1) install now copies to a same-directory "*.new.*" temp
  file and renames it into place, so a failed copy never reaches
  dest_plist at all; (2) the EXIT trap now restores pending_backup_dest
  from pending_backup_plist before removing the backup, rather than merely
  deleting it, covering whatever other failure might still reach it.
- New failure-injection test (the same differential shim) proves the fixed
  script's reinstall succeeds cleanly -- the injection never fires at all,
  since the vulnerable direct-to-destination copy no longer exists --  and,
  confirmed by stashing the pre-fix script, that the same injection there
  reproduces the exact truncate-and-delete-the-backup regression.

Full acceptance re-run: named macOS/launchd modules (115 tests) and the
bash-3.2/launchd-gated subset under a real compiled GNU bash 3.2.0 all pass;
bash -n and shellcheck clean on both scripts; scripts/validate.py and the
evidence manifest check pass. No macOS host ran any of this.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Re-register test_workflow_hardening.py after merging origin/main (#108)

scripts/validate.py flagged a hash/byte-count mismatch for
tests/test_workflow_hardening.py (part of #108's GitHub automation
closure) immediately after merging origin/main; re-registered via
host_receipts.register_file and re-normalized with evidence_manifest.py
--write. All other post-merge checks (component_matrix, new_host_grand_list,
evidence_manifest --check, validate_catalogs, validate_foundation,
build_verdicts --check, the full 2796-test suite) pass.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Round 3d: genuine recovery for both swaps (rollback, was_loaded/reload, atomic flip)

Codex's round-3 verification of 5310fe5 returned pass-with-findings for
everything except the two swaps' own recovery paths, which it reproduced by
direct failure and signal injection. Fixed by design (step tables added as
comments next to each swap; every row tested, including a failed rollback
and SIGTERM mid-swap sent via kill -TERM $PPID from a shim, matching the
coordinator's required method) rather than patched incrementally again:

bootstrap-macos.sh -- replaced the rename-aside-then-move-in swap with the
Homebrew Cellar/opt pattern: install into a versioned, never-reused
directory (tools/<id>-<version>-<stamp>); final_prefix becomes a SYMLINK,
flipped by a single python3 os.replace (genuinely atomic, unlike `ln -sfn`,
which is unlink-then-symlink). bin_dir's own symlinks are unchanged -- they
point into final_prefix/bin/* and transparently follow it through the
extra indirection, so they never need re-creating on a later flip. The one
non-atomic step this still has is a real, pre-3d final_prefix directory's
one-time migration aside (rename() cannot replace a directory with a
symlink in one call); the top-level cleanup() trap restores it if a
failure or a signal lands in that window, reporting the exact manual `mv`
if the restore itself also fails.

Also fixed along the way: final_prefix must never be canonicalized (it
silently resolved straight through the symlink to whatever it currently
targets, defeating the whole design -- caught by the first new failure-
injection test); pending_migration_prefix/_dest must be set BEFORE the
migration mv, not after (a signal in the gap between the mv completing and
the assignment left the trap unable to find what to restore -- caught by
the SIGTERM test); the test harness (_run_install_npm) needed the same
top-level cleanup()/EXIT trap the real script installs, which it had never
included, so failure/signal recovery was silently never exercised at all
until this was added.

launchd-agents.sh -- the EXIT trap now restores only by same-directory
rename (never `cp`, whose own failure was previously ignored before
deleting the backup unconditionally); deletes the backup only after that
rename succeeds, otherwise keeps it and reports the exact manual command;
records was_loaded (via pending_backup_reload_needed, set right after a
confirmed bootout) so restoring the backup from ANY exit path -- not just
the explicit bootstrap-failure branch -- also reloads the service, not
just its file. cmd_install's own explicit restore branch was removed
entirely in favor of the trap (round 3c/3b had two independent
implementations of the same recovery that had independently drifted
different bugs; one is now authoritative). The backup itself moved back
into dest_plist's own directory (reverting round 3b's N1 relocation under
state_dir): `mv` across a filesystem boundary silently falls back to a
non-atomic copy-then-unlink, so the restore-by-rename this fix depends on
is only genuinely atomic when the backup is guaranteed to be on the same
filesystem as dest_plist, which only its own directory can guarantee.

New failure/signal-injection tests (7 total: 3 for the atomic flip, 4 for
launchd's trap), each confirmed to fail against 96f1abe by stashing it.

Full acceptance re-run on this commit: full suite (2801 tests) passes
except the one evidence-registration lag scripts/validate.py itself always
flags immediately after an edit (resolved by the registration commit that
follows); named macOS/launchd modules (120 tests) and the bash-3.2/
launchd-gated subset (54 tests) under a real compiled GNU bash 3.2.0 all
pass; --plan clean under both bash versions; shellcheck and zizmor clean;
git diff --check clean. No macOS host ran any of this.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Round 3e: idempotent convergence for launchd, ownership-checked pruning for bootstrap

Codex's round-3e review of c18aea3 found two new Highs and four Mediums --
each earlier round's hand-written rollback had closed one edge case and
opened another (an unowned rm -rf, a partial hard-copy backup, a reload
skipped at a step boundary, cleanup left armed after success, a signal
racing a foreground mv, an unguarded rm aborting recovery under set -e).
Per the coordinator's explicit direction, this round changes the approach
instead of patching further: simplify toward idempotent convergence, the
way `brew services` works, with step tables and signal-injection tests
proving every row.

Both scripts now explicitly trap INT, TERM and HUP (`exit 130/143/129`
only; recovery logic stays in the EXIT handler alone) -- bash only
reliably defers a caught signal to a completed step when the signal is
explicitly trapped, not left at its default disposition (confirmed
directly: an untrapped SIGINT sent to a bash script with only an EXIT trap
registered does not interrupt it at all in this environment, while an
untrapped SIGTERM does -- the asymmetry Codex's finding 5 was pointing at).

bootstrap-macos.sh (High #1, unowned rm -rf): new prune_old_version()
deletes a superseded versioned directory only when its canonical form is a
direct child of the canonical tools/ directory with a name matching
<id>-<version>-<stamp>; an external target, a relative "../" target, a
non-matching name, or a symlink loop (canonical_path's python3
os.path.realpath handles cycles without hanging) is logged and left in
place, never deleted. Pruning is now genuinely best-effort: a rejected or
failed deletion warns and never fails the install. Four new tests, one per
adversarial target shape, confirmed to fail against c18aea3 (the function
does not exist there).

launchd-agents.sh -- full redesign, net roughly flat in size (+22 lines)
despite ~390 changed, since one pure function replaces four tracking
markers (pending_backup_plist/_dest/_reload_needed, pending_dest_tmp) and
their two independent, historically drifting recovery implementations:
- backup_plist is a hard link (`ln`, High #2), not `cp`: it either exists
  completely or not at all, so a partial copy can never overwrite the
  intact original. Same directory as dest_plist (round 3d), never this
  script's own state_dir.
- Staging (cp to dest_plist.new, then mv over dest_plist) is unchanged in
  shape but the name lost its PID suffix, matching the backup's own
  naming, so a stale artifact from a crashed run is visible to and
  reconciled by the very next run, not just the same process.
- reconcile_install replaces every marker: it reads dest_plist's identity
  against backup_plist (-ef, since the backup is a hard link) and
  launchd's own current loaded state, and converges to whichever ONE
  action that state calls for (drop the backup, reload, restore-and-
  reload, or -- if nothing converges -- keep the backup and name the exact
  re-run command). It is the ONLY recovery logic: called explicitly after
  a normal attempt, and identically as the EXIT trap for anything that
  interrupts one. `set +e` inside it, explicit checks throughout, per the
  coordinator's instruction that the handler must never itself become
  fatal under the script's own `set -e`.
- record_enabled_label now runs as soon as install actually commits to a
  label (not only on success), so a label that ends up needing attention
  is never re-refused as unowned on retry; a was_already_enabled snapshot,
  taken before that point, keeps the L2 bootout gate correctly checking
  ownership as it stood BEFORE this run, not after.

New tests: a shared _signal_shim helper plus a parameterized step x signal
matrix (ln and mv, TERM and INT) proving retry-converges for launchd; a
dedicated SIGTERM/SIGINT-parameterized migration test and three more for
bootstrap-macos.sh's prune_old_version. Every new/changed test confirmed
against c18aea3 by stashing the fix; several rewritten round-3b/c/d tests
(now asserting the new design's own honest behavior -- e.g. a fresh
install that never loads is left on disk with a manual-retry message,
never silently deleted) checked to still exercise the coordinator's
original findings under the new mechanism.

Full acceptance re-run on this commit: full suite (2821 tests) passes
except the one evidence-registration lag scripts/validate.py itself always
flags immediately after an edit (resolved by the registration commit that
follows); named macOS/launchd modules (126 tests) and the bash-3.2/
launchd-gated subset (56 tests) under a real compiled GNU bash 3.2.0 all
pass; --plan clean under both bash versions; shellcheck and zizmor clean;
git diff --check clean; guarded gitleaks clean. origin/main had not moved
since it was already merged into c18aea3's own parent, so no new merge was
needed this round. No macOS host ran any of this.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Round 3f: abort preflight on unresolved reconcile, path-verify launchctl state, poll through delayed teardown

Codex verified d51f062 FAIL with a High and two Medium regressions in
launchd-agents.sh's reconcile_install/cmd_install, plus an open delayed-
teardown boundary; bootstrap-macos.sh itself was confirmed sound and is
untouched here (its test matrix gained the same row/signal coverage).

- High: reconcile_install's own preflight failure is no longer ignored
  (`|| exit 1`, not `|| true`): a retry that still cannot converge now
  aborts before ever reaching the backup-overwriting `ln`, so the last
  known-good backup can no longer be silently destroyed and re-linked to
  unusable content.
- Medium (tri-state print): launchctl_label_state() reports loaded_here,
  loaded_elsewhere, not_found or unknown instead of a binary loaded check;
  unknown now takes no destructive action anywhere and keeps the backup.
- Medium (unrelated loaded label): ownership is recorded only inside
  reconcile_install, once a path match confirms dest_plist itself is what
  launchd has loaded -- never earlier -- and both reconcile_install's
  fresh-install path and cmd_install's own bootout gate refuse outright on
  loaded_elsewhere, on the first invocation and every retry.
- Delayed teardown: reconcile_install re-polls (reusing wait_until_unloaded)
  when this run's own bootout is still pending before trusting a "loaded"
  read, instead of declaring convergence with zero reload attempts.

Added four named regression tests reproducing Codex's three exact scenarios
plus the delayed-teardown case, each verified via differential (git stash)
to fail on d51f062 and pass here. Expanded the checked-in matrix to cover
every row of both step tables under failure, SIGTERM and SIGINT (was two
launchd rows and one bootstrap row), reusing and extending the existing
signal-shim helpers rather than one bespoke script per case.

Updated the in-script step table for the ownership-timing and preflight-
abort changes, and recorded the delayed-teardown re-poll as an untested
boundary for real-Mac acceptance in the platform notes: modelled against a
mock launchctl and real bash signal injection, never observed on native
launchd. Verified with the full adoption test suite (61 launchd + 78
bootstrap-macos tests) under both bash 5.2 and a real compiled bash 3.2.0,
shellcheck, zizmor, a manual bash-3.2 offline render/lint/install/status/
remove cycle, a manual bash-3.2 real-SIGTERM injection during install,
scripts/validate.py and friends, and guarded gitleaks (dir + git log range)
-- all clean. origin/main has not moved past this branch's base; no merge
needed. Evidence class unchanged: Linux-side local_integration/structural
plus real npm/Node/bash-3.2/signal behavior measured directly on this host;
no macOS host has run any of this.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Scope the pre-reinstall load-check test to cmd_install's own body

test_install_gates_its_pre_reinstall_load_check_on_ownership matched
cmd_remove's textually similar is_enabled_label + launchctl print gate
instead of cmd_install's round-3f one, so it passed for the wrong reason.

Added _shell_functions (the same top-level-function extraction helper
tests/test_adoption_bootstrap_macos.py already uses) to scope the
assertion to cmd_install's own body, and rewrote it to check the round-3f
gate directly: launchctl_label_state is called before any bootout, and the
loaded_elsewhere and unknown case arms both refuse (exit 1, no bootout)
rather than proceed.

Verified the new test actually exercises cmd_install and not cmd_remove by
temporarily replacing cmd_install's gate with the old pre-round-3f binary
launchctl-print check (no path verification, no loaded_elsewhere/unknown
refusal) and confirming the test fails against that mutation, then
restoring the gate byte-for-byte (diff against HEAD is empty for
launchd-agents.sh; only the test file and evidence manifest changed here).

Re-ran the full test_adoption_launchd suite (61/61) under bash 5.2 and a
real compiled bash 3.2.0, plus bash -n on both interpreters, shellcheck,
scripts/validate.py, evidence_manifest.py --check (re-registered via
host_receipts.register_file) and guarded gitleaks dir -- all clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Round 3g: replace launchd's transactional backup/reconcile with brew-services semantics

Codex's review of 0e1e0d4 found four new Medium regressions in the round
3b-3f backup/reconcile design (remove booting out an unrelated service on
stale state; a print error on fresh install deleting an already-loaded
plist; a signal during bootout or restore leaving the service stopped with
no retry built in; an ownership-write failure falsely reporting success) --
the fifth straight round to find a new recovery defect in that machinery.
Rather than patch a sixth, launchd-agents.sh now converges on the
maintained upstream pattern: Homebrew/brew Library/Homebrew/services/
cli.rb @ 8e3a5dc0a7 (2026-09-07), which has no backup, no rollback and no
ownership file, and recovers by re-running the command.

- Removed record_enabled_label, is_enabled_label, forget_enabled_label,
  the enabled-labels state file, reconcile_install, was_already_enabled
  and the hard-linked backup entirely.
- Stateless ownership: "ours" means the label is one of this script's own
  AND, if it is loaded at all, launchctl print's own "path =" line names
  this script's destination plist -- never a stored record. install and
  remove share the exact same launchctl_label_state check (tri-state plus
  elsewhere, unchanged from round 3f) and both refuse outright, touching
  nothing, on loaded_elsewhere or unknown.
- install: stage a temp file beside the destination and lint it; if
  loaded_here, bootout and wait until confirmed unloaded (refusing
  otherwise); rename into place; launchctl enable + bootstrap; on
  bootstrap failure, leave the new plist in place (no rollback -- there is
  nothing to roll back to) and print the exact recovery commands.
- remove: on loaded_here, bootout, wait, then delete; on not_found, delete
  directly if a file is present; refuse on loaded_elsewhere or unknown.
- The EXIT trap now only removes the current run's own still-inert temp
  file; INT/TERM/HUP stay explicitly trapped so bash always defers to a
  completed step.

Replaced every round 3b-3f backup/reconcile test with a convergence matrix
(cp/bootout/mv/bootstrap x {failure, TERM, INT}, then a fault-free install
retry converging to NEW loaded, then a fault-free remove leaving nothing
loaded and no file) plus named tests for loaded_elsewhere refused by both
install and remove, unknown touching nothing on both, and Codex's exact
remove reproduction (a previously installed label's plist on disk, its
label now loaded from a different, unrelated path). All new/changed tests
verified via differential (git stash on the script alone) to fail on
0e1e0d4 (19 of 55 tests fail there, the rest being unaffected render/
template/structural checks); all 55 pass here.

adoption/platforms/macos-arm64.md carries the dated decision (chosen: this
design, with the Homebrew citation; rejected: the transactional design,
because rounds 3b-3f each found a new recovery defect in it; would
overturn: a real-Mac observation where a re-run does not converge, or a
requirement to preserve a locally edited plist) and keeps the untested
delayed-teardown-polling boundary note, updated for the new call sites.

launchd-agents.sh shrank from 758 to 593 lines (38310 to 28554 bytes)
against 0e1e0d4, removing the transactional machinery outright rather than
adding another layer of recovery logic on top of it.

Merged origin/main (bb5228d, unrelated: release-tag pin bump) -- no
conflicts with this file set. Verified: full test_adoption_launchd (55/55)
and test_adoption_bootstrap_macos (78/78, unaffected) under bash 5.2.21
and a real compiled bash 3.2.0; a manual bash-3.2 offline render/install/
status/remove cycle and a manual bash-3.2 real-SIGTERM injection during
the rename step; shellcheck, bash -n on both interpreters, zizmor,
git diff --check, scripts/validate.py and friends (evidence re-registered
via host_receipts.register_file), and guarded gitleaks (dir) -- all clean.
Evidence class unchanged: Linux-side local_integration/structural plus
real bash-3.2/signal behavior measured directly on this host; no macOS
host has run any of this.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Round 3h: pin and download the llama-embed model, add the embedding acceptance check

A 2026-09-23 readiness audit found two defects in PR #94:

1. The llama-embed launchd agent's template ran llama-server with no
   model argument at all -- KeepAlive would restart it in a loop forever
   on a real Mac, never actually serving embeddings. Fixed:
   - com.native-stack.llama-embed.plist.template now passes -m
     ${EMBED_MODEL_PATH}.
   - adoption/pins-macos-arm64.json gained a separate top-level "models"
     section (never "tools", so it is never swept into
     adoption/manifest.json's protected profile-component coverage) for
     ggml-org/embeddinggemma-300M-GGUF @ 0f741b5a6585bd53aeb15cd1372c56f2a0f65e12,
     embeddinggemma-300M-Q8_0.gguf, sha256-pinned, with its own
     huggingface_lfs_oid_plus_local_rehash checksum source.
   - adoption/bootstrap-macos.sh gained install_embed_model, a dedicated
     download-and-verify step (reusing fetch(), the same helper every
     other pin already uses) run unconditionally after the tools[]
     loop, fail-closed on a null or mismatched sha256, into
     $ECO_INSTALL_ROOT/state/models/.
   - launchd-agents.sh's cmd_render and adoption/hosts/macos-example.json
     both carry EMBED_MODEL_PATH now, alongside the existing HOME/
     ECO_ROOT/AI_MEMORY_URL host values.
   - New tests: the rendered plist's -m argument matches EMBED_MODEL_PATH
     (test_adoption_launchd.py); install_embed_model, extracted verbatim
     and run against a real (curl-shimmed) fetch(), installs on a
     matching digest and fails closed on a mismatch or a null sha256
     (test_adoption_bootstrap_macos.py); the models[] pin's own shape,
     URL-matches-revision and separation from tools[].

2. No Linux reference vector existed to compare a macOS response
   against. Fixed:
   - evidence/artifacts/macos-embed-reference-20260923/ carries the
     captured reference (llama.cpp b11057 ubuntu-x64 CPU build, the
     identical pinned GGUF, 768 dimensions, L2-normalized, repeat cosine
     1.0, full request/response provenance), registered via
     host_receipts.register_file.
   - tools/adoption/embed_acceptance.py (stdlib only) sends the
     reference's exact request body to a running llama-server, checks
     dimension and cosine >= threshold against the reference embedding,
     and prints a JSON result; understands both the OpenAI-compatible
     and llama.cpp-native response shapes, and reports a network or
     malformed-response failure as JSON, never a bare traceback. Six new
     tests (a local stdlib HTTP server stands in for llama-server) cover
     an exact match, both response shapes, a dimension mismatch, a
     below-threshold cosine, and a connection failure.
   - .github/workflows/adoption-bootstrap.yml's bootstrap-macos job now
     caches the pinned model (actions/cache, keyed on its sha256, not a
     version tag), installs and health-checks the llama-embed launchd
     agent, runs embed_acceptance.py against it, uploads the JSON result
     and the startup log, and best-effort-greps the log for which
     backend it reports -- labelled explicitly as hosted-runner evidence
     only (a GitHub-hosted macOS runner is CPU-only, never a claim of
     Metal actually engaging; a real Mac run is what would show that).
   - adoption/platforms/macos-arm64.md's "What a hosted run proves" and
     "Acceptance test for the 24 GB default" sections both updated: the
     reference vector now exists, and the one command to run it is
     documented, still bounded to "unrun on any Mac, hosted or
     otherwise" for the acceptance itself.

Verified: full test_adoption_bootstrap_macos (93/93) and
test_adoption_launchd (56/56) under bash 5.2.21 and a real compiled bash
3.2.0; bash -n and shellcheck clean on both scripts; zizmor clean on the
workflow; git diff --check; scripts/validate.py and friends (evidence
re-registered via host_receipts.register_file); guarded gitleaks dir --
all clean. Evidence class unchanged and explicit throughout: Linux-side
local_integration/structural, a real captured Linux reference vector, and
(once this round's hosted CI step runs) CPU-only hosted-runner evidence;
no macOS workstation has run any of this.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Round 3i: canonical-path comparison, EINPROGRESS handling, cache-before-verify, embed_acceptance transport errors, macOS recording-tooling CI

Codex's review of e9334e9 returned FAIL with three Medium findings and one
Low, all with narrow, well-isolated fixes:

1. launchd-agents.sh's loaded-state classifier used `-ef` (device+inode
   identity), so a DISTINCT hard-linked filename sharing dest_plist's own
   inode incorrectly compared loaded_here, letting remove bootout and
   delete a plist for a label actually loaded from an unrelated,
   hard-linked path. Replaced with canonical_plist_path: resolves symlinks
   in a path's parent directory only (cd -P, falling back to python3's
   os.path.realpath), then reattaches the final component unresolved --
   neither mechanism ever collapses a hard link's own distinct pathname
   the way inode identity did, since a hard link has no stored "canonical
   name" to resolve to.

2. A `launchctl bootout` returning EINPROGRESS (teardown already under
   way, from this run's own retry of an earlier interrupted attempt, or
   another process) was treated as an unconditional failure, abandoning
   an in-progress teardown with no wait and no reload ever attempted. The
   pinned Homebrew cli.rb @ 8e3a5dc0a7 (services/cli.rb, ~line 311)
   retries on exactly Errno::EINPROGRESS (Darwin 36) for the identical
   reason. New shared bootout_and_wait absorbs it the same way, then
   polls (reusing wait_until_unloaded) rather than refusing outright.
   Self-discovered along the way: the first implementation captured its
   result via `result="$(bootout_and_wait ...)"`, which forks a subshell
   -- empirically verified that a signal delivered while a foreign
   command runs INSIDE that subshell is silently swallowed there and
   never reaches this script's own INT/TERM/HUP traps at all. Fixed by
   using a plain function call plus a global result variable instead,
   keeping bootout_and_wait in the same process as its caller.

3. adoption-bootstrap.yml's embedding-model cache restore ran AFTER
   bootstrap's own checksum verification and used a literal hardcoded
   sha256 as its key. A restore running that late could silently
   overwrite the just-verified model with stale cached bytes, with
   nothing left to re-check them, and a changed pin would keep matching
   the same stale cache entry. Fixed: a new step reads the pin's sha256
   into a step output, the cache restore (now keyed on that output) moves
   to before the bootstrap step, and bootstrap-macos.sh's own fetch()
   (unchanged, already correct) transparently re-downloads and
   re-verifies any mismatched destination on the exact same path a
   genuine cache miss takes.

4. embed_acceptance.py caught urllib.error.URLError around the request,
   but response.read() can raise a bare TimeoutError (a stall mid-body,
   past urlopen's own connect-phase timeout) or http.client.
   IncompleteRead (the server closing before its declared Content-Length
   was satisfied) -- neither caught, leaving stdout empty and the process
   exiting via an uncaught traceback instead of the documented JSON
   failure contract. Both now caught explicitly, plus a catch-all OSError.

Also, this round (peer-update-audit gap adoption_macos, medium): the
catalog's own recording and verdict scripts (host_receipts.py,
component_matrix.py, new_host_grand_list.py, build_verdicts.py,
validate_convergence.py, release_due.py) had never run on macOS CI or
against macOS's own system Python -- every existing gating job for them
runs on ubuntu-24.04 only. adoption-bootstrap.yml's validate-macos job now
gates on all six, then runs a recording smoke: installs the profile,
records a real host_receipts.py record receipt (component codex,
--from-stack-commands, --evidence-class native_proven -- hosted-runner
evidence only, never a workstation acceptance receipt) inside a throwaway
copy of the checkout ($RUNNER_TEMP/rec, never the real one, so nothing is
committed), then re-validates and re-checks the matrix/grand-list in that
same copy. Runs once against the manifest-pinned Python line and once
against the runner's own system /usr/bin/python3 if it meets the newly
declared Python 3.9 floor (adoption/bootstrap.md, adoption/platforms/
macos-arm64.md's new "Recording and verdict scripts" section), skipping
with a message otherwise.

Added named tests for all five changes (a hard-linked alias never
compares loaded_here; a retry converges while an earlier run's teardown
is still asynchronously in progress, fixing the bootout-signal test shim
Codex flagged for clearing the marker before signalling and so hiding the
async case; CI step order and the derived cache key; two embed_acceptance
transport-failure modes over a raw-socket server; structural checks on
the new validate-macos steps), each verified via differential (git stash)
to fail on e9334e9 or the pre-fix file. Full suites: test_adoption_launchd
59/59, test_adoption_bootstrap_macos 107/107.

Verified: both scripts under bash 5.2.21 and a real compiled bash 3.2.0
(syntax and shellcheck), PLUS two direct, real-SIGTERM/real-shim
reproductions under bash 3.2 outside the Python harness (the hard-link
refusal, and the EINPROGRESS retry actually converging); zizmor and
actionlint both clean on the workflow; git diff --check; scripts/
validate.py and friends including host_receipts.py validate and
build_verdicts.py --check (evidence re-registered via host_receipts.
register_file); guarded gitleaks dir -- all clean. origin/main has not
moved past this branch's base; no merge needed. Evidence class unchanged:
Linux-side local_integration/structural plus real bash-3.2/signal
behavior measured directly on this host; the new CI steps are themselves
unrun until a real macos-15 job executes them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Fetch full history in validate-macos so host receipt catalog revisions resolve

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Resolve symlinked plist leaves in canonical_plist_path; make the incomplete-body test drain the request

canonical_plist_path now follows a symlinked leaf (bounded at 40 hops) so a
plist launchctl reports by its resolved target compares loaded_here, while a
hard link stays a distinct path (Codex round-3i finding 1). The raw-socket test
server drains the whole request and half-closes, so the client sees
IncompleteRead rather than an occasional ECONNRESET (30/30 runs pass).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Skip the PyYAML structural workflow tests where PyYAML is absent (repository convention)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Give the macOS recording smoke a per-run --identity (required since #117)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Record the 2026-09-23 hosted macOS smoke receipt (launchd, embedding acceptance, recording smoke)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Round 3j: fix 8 unresolved Codex-connector review threads blocking the #94 merge

1. launchd-agents.sh:77 -- llama-embed was left out of the default install/
   status/remove label set, even though bootstrap now downloads and
   sha256-verifies its model unconditionally (round 3h). Added it to
   default_labels; --label was the only way to reach it before this.

2. launchd-agents.sh:526 -- install created the plist's own declared
   StandardOutPath/StandardErrorPath/WorkingDirectory BEFORE the ownership
   check, so a refused install (loaded_elsewhere or unknown, which must
   touch nothing) still left new, empty directories behind. Moved
   directory creation to after the ownership check's case statement.

3. macos-arm64.md:217 -- "What a hosted run proves" was stale (cited a
   2026-09-22 run, claimed launchd and the embedding acceptance were both
   still unrun). Rewritten for run 35875188590 at head 75a6e0d: launchd
   genuinely bootstrapped and booted out both qdrant and llama-embed;
   qdrant health reported 1.19.1; the embedding acceptance passed with
   cosine 0.99944 and the startup log reports Metal device MTL0 engaging
   for the layers it covers; the recording smoke passed on both
   actions/setup-python 3.13.15 and the runner's own system
   /usr/bin/python3 3.9.6. Kept the hosted-runner-only scope throughout --
   no platform_status change, no workstation acceptance claim.

4. qdrant.plist.template:28 (the most important one) -- nothing on a
   clean host ever created $ECO_ROOT/config/qdrant.yaml; only the hosted
   CI step's own disposable file did. Added provision_qdrant_config to
   bootstrap-macos.sh: mirrors examples/qdrant.yaml.example exactly
   (storage under $ECO_ROOT/state/qdrant, loopback, port 16333 -- the
   same port adoption/hosts/macos-example.json's QDRANT_URL and
   recipes/README.md both name), plan-mode aware, and never overwrites an
   existing file. The hosted launchd step no longer writes its own
   disposable config or uses port 6333; it uses the provisioned one at
   port 16333, exactly what a real install would use.

5. ai-memory.plist.template:30 -- added --workspace local --project
   native-agent-stack, matching the selected native recipe's own serve
   invocation byte for byte (recipes/README.md:244; the same scope
   docs/foundation-stack.md:51-52 and every other documented ai-memory
   invocation in this project already use).

6. adoption-bootstrap.yml:13 -- the paths triggers covered adoption/**
   and tools/adoption/** but missed most of what validate-macos's
   recording-tooling gate and recording smoke actually consume. Added the
   embed reference artifact, host_receipts.py, component_matrix.py,
   new_host_grand_list.py, build_verdicts.py, platform_status.py,
   validate_convergence.py, release_due.py and manifests/evidence.json to
   both push and pull_request path lists.

7. launchd-agents.sh:517 -- lint only checked plist syntax, never that a
   rendered plist's own Label key matches its filename (also the label
   every other subcommand addresses it by, and what launchd itself
   expects to agree). Checked once syntax lint has already passed, via
   pure bash parameter expansion (no new basename/dirname dependency).

8. launchd-agents.sh:620 -- status always exited 0, even when every
   requested label's launchctl print failed (absent, 113, or launchd
   itself unavailable, any other nonzero exit) -- a caller checking only
   the exit code never saw either case. Now returns non-zero whenever any
   requested label's print fails…
seathatflowsinourveins added a commit that referenced this pull request Sep 24, 2026
…er, no duplicate backtrader gap

- L1: a sentence dropped between two kept ones is marked [...] so a joined excerpt never reads as
  contiguous upstream text.
- L2: counts and superlatives outside the word-boundary filter ("1B+", "#1", "most complete",
  "leaderboard") are dropped.
- L3: the backtrader gap is removed; its home layer already carries the beyond lane's
  unmaintained_signal.
Pinned commits are unchanged (reused from the reviewed receipts); 10 receipts change in text only.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 24, 2026
…eviews for its newcomers (#153)

* Add the 2026-09-23 landscape-sweep lane to manifest-20260923: 98 proposals, 11 survivors with source reviews

A per-layer sweep of all 32 layers (one researcher per layer, a facts/identity and a fit/standing
refuter per layer, every worker at effort max) proposed 98 repositories the catalog did not know;
11 withstood both refuters. The manifest is rebuilt with the unchanged generator from the original
private inputs plus this lane (the original inputs alone reproduce the committed manifest byte for
byte). Survivors carry a neutral source_review file at a pinned upstream commit, registered and
listed first in their evidence, for the blind re-record lanes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Close the #153 review: catalog-recorded repos known, upstream-only receipts, existing newcomers covered

- F2: repositories already recorded in catalogs/foundation/automation.json (active osv-scanner and
  Scorecard lanes; dated considered_not_activated decisions for Renovate and gh-aw; harden-runner)
  are known, not candidates; the two active ones become open gaps on their layers instead.
- F1/F3/F4: receipts carry only the upstream's own words (repository description and verbatim
  README excerpts at the pinned commit) and only the pinned tree/README URLs; popularity, standing,
  recency, status and comparison sentences are filtered; winner names match whole words; nav-only
  excerpts and user/sponsor sections are skipped.
- agent-lab-17's request: the manifest's 57 existing surviving newcomers (51 repositories, 6 of them
  Hugging Face models) get the same neutral source-review files, prepended to their evidence.
- gitleaks: a rule-scoped, exact-file, whole-line allowlist for the manifest's "pin" git commit ids,
  which the sourcegraph-access-token rule's keyword activated once lane text named Sourcegraph; with
  regression tests that other fields and other files stay detected.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Record two verified hard signals as open gaps: poppler's catalog URL is a mirror; backtrader unpushed since 2024-08

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Close the round-2 lows of #153: marked excerpt splices, standing filter, no duplicate backtrader gap

- L1: a sentence dropped between two kept ones is marked [...] so a joined excerpt never reads as
  contiguous upstream text.
- L2: counts and superlatives outside the word-boundary filter ("1B+", "#1", "most complete",
  "leaderboard") are dropped.
- L3: the backtrader gap is removed; its home layer already carries the beyond lane's
  unmaintained_signal.
Pinned commits are unchanged (reused from the reviewed receipts); 10 receipts change in text only.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Close Codex's #153 review and merge the critic follow-up round

- Codex P1 (RD-Agent): its comparison scores on a final segment the feedback loop never sees.
- Codex P2 (TruffleHog, 2 layers): the verification arm uses a controlled, revocable credential.
- Codex P2 (MinerU): stale v1.0-era OmniDocBench figures and CJK framing removed.
- Codex P2 (Marker): model-weight license terms recorded separately from the code license.
- Codex P1 (stopped run): the first run's 9 discovery returns are retained, label-free, with call
  counts and whether the completed run re-proposed each repository; its usage is unknown.
- The completeness critic's follow-up round (same two-refuter rule) is merged: 107 proposals,
  11 survivors (adds PaddleOCR, ollama, betterleaks, claude-code-action), with prior documentary
  records disclosed; completed-run usage recorded (119 agents, 12.1M subagent tokens).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Close the final #153 review: PaddleOCR rank corrected, claude-code-action records disclosed

- M1: PaddleOCR-VL-1.6 is 3rd on OmniDocBench v1.6 (README at f133a71e9e), not 1st; its proposed
  label is lowered to keep_but_compare because its fit vote called the rank-1 basis void.
- M2: claude-code-action's prior records are cited correctly (targeted_candidate with an executed
  CI smoke arm in the 2026-09-22 SDK sweep; decision HOST-09) in the lane limits and its comparison.
- L1: RD-Agent keeps its reproduce-the-published-advantage stay condition.
- L2/L3: a stale Release claim is corrected and a private scratch name is removed from evidence.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Word the survivors' prior-records limit precisely: no adopt or reject decision (HOST-09 is keep-but-compare)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: seathatflowsinourveins <234074349+seathatflowsinourveins@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins deleted the codex/mac-worker-timeout-current branch September 25, 2026 18:47
seathatflowsinourveins pushed a commit that referenced this pull request Sep 29, 2026
…ere-documents are data, ps -fu Eve passes

Failed first (tests/test_secret_path_guard.py, run against the guard of 2f097bf): 64 subtest
failures on the new rows, none of them passing by luck. Blocked before and now allowed or
blocked as they should be: 11 BLOCKED rows returned None ("# macOS" then ps -E, "# check ..."
then systemctl --user show-environment, "# run it" then systemd-run --pipe --wait printenv,
echo ${#PATH}; printenv, echo $#; printenv, gh api .../issues/1#c; cat "$PAPER_ENV_FILE",
echo a#b; printenv, two commands with a trailing comment, echo "$(date # )" whose printenv line
follows the comment, a comment holding quotes and a backquote, and a here-document heading
followed by printenv), plus 22 behind rtk proxy and 11 behind a keyring exec. 8 ALLOWED rows
returned a reason: echo ok # "$(printenv)", printf '%s' $'it\'s "$("printenv")"', ps -fu Eve
(two rows) and four commit-message patterns with a quoted here-document. 11 scanner-table rows
and 1 expected-pass-through row (a shell reading a quoted here-document in a substitution)
differed, and all 10 subtests of the five real commit messages (cf58426, 32a6cd2, a84fa7c,
88af2ba, b553826, verbatim, in the standard git commit -m "$(cat <<'EOF' ...)" and the gh pr
create --body forms) failed. The bug behind the first group is older than this work: tokenize()
joins lines with ";" and shlex reads a "#" anywhere, "$#" and "a#b" included, as a comment, so a
"#" dropped the whole rest of the command (found by the 2026-09-29 review).

Design: scan_shell() (substitution_bodies() now calls it) returns the bodies and the comment
spans from the same single pass. A comment is a "#" that starts a word (at the start of the text
or after a blank or one of ;&|()<>, and not right after a quote, escape or substitution), outside
quotes and outside a double-quoted substitution, to the end of the line; inside a $( ) body it
also runs to the end of its line, so a ")" in it closes nothing, and in a backquote body it ends
at the closing backquote. tokenize() removes the spans and sets shlex's commenters to "", and
command_segments() also reads the command as shlex did before (for commands up to 200,000
characters) and adds only the segments that reading lacks, so nothing the guard read before is
dropped. In a here-document body a comment hides the rest of that body, as before, not the text
after the terminator, so a script written through a here-document (its first line is #!) stays as
unread as it was. $'...' is data up to the first quote that no backslash escapes, except inside
double quotes. A here-document whose delimiter is quoted (also partly: <<'E'OF) is data: the body
that scan_shell returns loses it, keeping the operator line and the terminator; an unquoted
delimiter keeps its body, as at the top level. ps_shows_environment() skips the word after a
cluster that ends in a value option (ps -fu Eve). The three vacuous BLOCKED rows are marked as
regression rows and two are replaced by rows that pass at base, echo "$(echo \"$(printenv)\")"
and echo "$(base64 ~/.ssh/id_ed25519)".

Checked: oracle 65 cases, 0 mismatches; suites: only test_host_profile_copy_is_verbatim fails
(139 tests). Differential against base c26800f over 54,126 generated commands: 0 loosened, 0
unexplained reason changes, 0 newly blocked without a new form. Random grammar fuzz against
base, 160,000 strings in 4 seeds: 0 loosened, 0 exceptions. Real repository, by realistic shape:
785 fenced blocks, 4,338 fenced lines and 201 scripts written through a quoted here-document: 0
newly blocked; the synthetic shape of a whole script passed as one command string newly blocks 4
of 201, none of them a true positive and all long-standing readings that the old "#" had
hidden (three array literals x=(env -i ...) and a python set(...)). Commit messages: the
standard pattern newly refuses 0 of 842 (both guards refuse the same 8; the previous tip newly
refused 3 of the 842 and 6 of 1,865 across all refs); the top-level
git commit -F - shape newly blocks the two messages of this work that quote the new ps -E and
systemd-run rules in prose, by those rules and not by the comment change. 1 MB timing: a 840 KB
git commit -m "$(cat <<'EOF' ...)" takes 8.5 s at base and 8.9 s now (two readings would take
17.8 s, so commands above 200,000 characters are read once).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 29, 2026
…er-bash parser (#516)

* Add failing-first N1 probes and must-stay controls for M4

N1 is the #432 fixup3 residual: openers() does not track "$(" inside
double quotes, and a quoted string that a shell runs is appended
without its own heredoc resolution, so heredoc data reads as executed.

Adds three tests through the existing measure/exports bridges on all
three shell carriers plus fetchKind:
- p1 (git commit -m "$(cat <<'EOF' ...)" with a line-start curl), its
  body variants with ', " and ( ), p2 (gh api in the body), p3
  (bash -c 'cat <<EOF > x.sh ...') and a double-quoted p3 variant;
- must-stay controls: an executed curl in "$( )", a shell heredoc in
  "$( )", an escaped substitution in a double-quoted run string, the
  command after a heredoc in "$( )" (guards against resolving one
  heredoc twice), the outer-shell heredoc in a double-quoted run
  string, unquoted $( ) and the "Subject (scope)" body;
- the nested-quote case from design 1.5.

At cf3fb72e the new tests fail 23 subtests (p1, the ' variant, p3 and
the double-quoted p3: shell fetch 1 != 0 and fetchKind 'fetch'; p2:
unclassifiable 1 != 0; nested quotes: 0 != 1 and fetchKind None); the
must-stay test passes. Source: POSIX.1-2024 XCU 2.6.3 Command
Substitution, checked with GNU bash 5.2.21 and dash.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add failing-first N1 tests for strings a shell runs

The N1 fix analyzes a string that a shell runs (RUN_QUOTED) as that
shell's input in both modes, instead of appending it verbatim unless
inlineHttp is set. Two consequences get their own tests first:
- in a double-quoted run string the outer shell removes the backslash
  before " (POSIX.1-2024 XCU 2.2.3), so `bash -c "echo \"a; curl
  ...\""` runs only echo: base reads a confirmed shell fetch;
- a string the inner shell runs in turn is analyzed too, so the curl
  in `ssh host 'bash -c "curl ..."'` is confirmed: base counts it in
  neither the confirmed nor the unconfirmed fetches, so the gate's
  lower bound could not see it.

Against the cf3fb72e kernel the test fails 7 subtests (1 != 0 and
0 != 1 on each carrier, fetchKind ['fetch', None]). GNU bash 5.2.21
and dash print `a; echo RAN` for the escaped form.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Resolve heredocs inside "$( )" and in strings a shell runs (N1)

Fixes the #432 fixup3 residual N1 in child-usage.mjs. POSIX.1-2024 XCU
2.6.3: inside double quotes the text from "$(" to the matching ")" is
tokenized recursively; 2.2.3: in a double-quoted word the outer shell
removes the backslash before $ ` " \ and newline. GNU bash 5.2.21 and
dash agree on every shape below.

- step() is one frame machine (', ", `, and $( with depth) shared by
  phase 1 openers() and the phase 2 closeQuote()/matchParen(), so the
  phases agree on where a quote or substitution ends. With no frame
  open it reads exactly as the old openers(); only a double-quoted
  word gains $( and backquote frames, so a << inside "$( )" is found
  and its body resolved like any other heredoc.
- executedTrace splits into resolveHeredocs (phase 1) and scanQuotes
  (phase 2). The naive end-of-quote scan becomes closeQuote; a "$( )"
  body in double-quoted data is spliced as "$(" + scanQuotes(body) +
  ")" with exact offsets (no second phase 1 over it).
- A string a shell runs (RUN_QUOTED) is analyzed with the full
  executedTrace in both modes, a double-quoted one after unquoted()
  removes the outer shell's escapes at its own level. A shared set of
  resolved operator offsets keeps a heredoc the outer shell resolved
  in "$( )" from reading later lines as its body again.
- NESTING_LIMIT (32) bounds recursion: deeper analyses read as data.
  Without it '"$(' x 3000 and a 3000-deep shell heredoc chain throw
  RangeError, which would abort a sweep.
- The paren-free DOUBLE_QUOTED_DATA regex is removed; the scanners are
  linear character scanners (no new regex; CodeQL js/redos).

The failing-first tests from 1d206a12 and 54299c44 now pass; the
must-stay controls are unchanged. SHA256SUMS is regenerated in a
later stage.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Document the N1 reading and add kernel unit cases

The workflows README replaces the N1 limit paragraph with the rule the
kernel now follows: "$( )" inside double quotes is read recursively
(POSIX.1-2024 XCU 2.6.3), a string a shell runs is analyzed as that
shell's input after the outer escapes are removed (XCU 2.2.3), and
reading text as data moves a raw detector match to the unconfirmed
count instead of dropping it. It states the remaining limits:
backquoted spans inside double quotes, case patterns inside "$( )",
the 32-level nesting bound, and text bash rejects as incomplete.

tools/skill-usage/README.md notes that the legacy Python executed_text
rule predates this reading and can disagree on these shapes.

test-child-usage.mjs adds fetchKind and exact executedText cases for
the N1 shapes and a nesting witness. Against the cf3fb72e kernel the
three new cases fail (95/98); with NESTING_LIMIT removed the witness
fails (97/98, RangeError); with the fix all 98 pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* State how the N1 reading moves the M4 lower bound

The previous wording said reading text as data leaves the denominator
of routed_share_lower_bound unchanged. That holds for raw detector
matches only. Measured against the cf3fb72e kernel:
- `echo "$(echo "a" && curl http://127.0.0.1:9/)"` moves from
  unconfirmed to confirmed loopback, which leaves the denominator;
- `echo "$(\curl ...)"` gains a confirmed fetch no raw detector sees;
- `echo "$(echo " ; \curl ...")"` loses a confirmation the old
  reading created (bash prints it as data), so the bound can rise.
The README now states those three cases. It also adds a limit: a
backquoted span outside quotes is still read as plain command text,
so a quote it leaves open runs past its closing backquote.

test-envelope (254/254), test-usage-receipts (18/18) and
test-contract-mutations (74/74) pass; they read this README as
routing_doc and contract_doc.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add failing-first CLI-lane, state and proxy-position fixtures

Add the U1 design section 7 fixtures before the kernel change. The
Python integration cases feed measureTranscript, aggregateMeasurements
and the sweep CLI: the stack.json lane commands (:1003, :797, :579, :762,
:519, :629, :280, :407), wrappers, compound commands and substitutions,
runners, every mcporter call form, negatives (data, lookups,
registrations, near names, gcm and serena-hooks), version and help,
remote and unresolved programs, the call states, the rtk proxy
population by command position with prefix_rule_calls, M3 and M4
by_carrier, aggregation with actors_with_success, and ID-free sweep
output. The node suite adds commandInvocations unit cases.

Sources: #381 AA-PLAN PR-A item 3 and M6/M14; POSIX.1-2024 XCU 2.9.1-2.9.4
and 2.6.3; rtk-ai/rtk@1d87b8e7 src/main.rs:68-90,708-713,3008-3042 and
the installed rtk 0.50.0 (proxy -- and --skip-env bind, -v after proxy
is the program); openclaw/mcporter@93e0916c src/cli.ts:115-364,
src/cli/command-inference.ts:10-95, call-arguments.ts:79-233 and
call-command.ts:114-173,309-348 (a first positional with = is the
selector). Before the fix: 10 Python tests fail (102 subtests) and the
8 new node cases fail on the missing export.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Count CLI lanes and rtk proxy calls by command position

Export commandInvocations: every simple command of the executed text
(POSIX.1-2024 XCU 2.9.1-2.9.4, 2.6.3) is resolved past reserved words,
assignments, wrappers (timeout, env, nice, stdbuf, command, exec, time,
nohup, sudo 1.9.15p5, xargs), package runners and rtk proxy to a lane
executable by exact basename. mcporter operations and a call's
downstream server follow openclaw/mcporter@93e0916c (cli.ts,
command-inference.ts, call-arguments.ts, call-command.ts); HTTP and
stdio selectors read (http) and (stdio), and no host is emitted. rtk
proxy follows rtk-ai/rtk@1d87b8e7 src/main.rs:68-90,708-713,3008-3042
and the installed rtk 0.50.0 (no shell runs the proxied program).

One memoized reading per call drives carrierOf (M3 and M4 by_carrier),
measurement.proxy (adds invocations, nested, in_ctx_code and
prefix_rule_calls) and the new measurement.cli_lanes with per-call
states (AA-PLAN "Successful" and M14; not_executed from observed
transcript markers). aggregateMeasurements sums the fields and adds
actors_with_success; the sweep limits gain the static limits.
executedTrace takes an optional marks collector for ssh-run text and
data-heredoc substitutions; its output is unchanged (0 differences in
executedText and fetchKind over 24,489 inputs).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Read mcporter URL and stdio selector forms with char scanners

Replace the character-class regexes in httpUrl and httpToolSelector,
the path-prefix alternation for an ad-hoc stdio selector and the
blank split of an inline npx server with character checks and the
existing shellWords splitter. The brief asks for linear character
scanners rather than regexes in this kernel (CodeQL js/redos flagged it
before). Behaviour is unchanged: the rules still follow
openclaw/mcporter@93e0916c src/cli/http-utils.ts:1-69,
call-argument-values.ts:68-83 and ephemeral-target.ts:134-158, and the
node suite passes unchanged (107/107).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Document CLI lanes, call states and the command-position proxy rule

Add a CLI-lanes subsection to the PR-A measurement fields: the lane
executables with the gcm and serena-hooks exclusions, the resolver
(wrappers, runners, rtk proxy), each cli_lanes field, the mcporter
operation and downstream-server precedence of mcporter v0.14.1, the
call states with the observed not-executed markers and their M14
mapping, the static limits, and the changed meanings of proxy.calls,
M3 by_carrier.rtk_proxy and m4.by_carrier with prefix_rule_calls kept
for #369-era comparison. Record the decision that proxy.in_ctx_code
stays outside M6 (Gate A plan M6 and rtk rows) and state that
rtk_parts, proxy_parts and model_typed keep their prefix rules. The
skill-usage README points Codex readers to the same fields and states
that declined items and legacy exec headers still read succeeded until
the adapter maps them.

Sources: rtk-ai/rtk v0.50.0 src/main.rs:68-90,3008-3042; openclaw/mcporter
v0.14.1 src/cli and docs/call-heuristic.md; mksglu/context-mode v1.0.169
src/exit-classify.ts:15-33; POSIX.1-2024 XCU 2.9.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add a failing case: exec -- reads as an unresolved program

U1 design 2.2 reads only the POSIX exec form and marks any word that
starts with '-' after exec as unresolved, because bash's exec options
are not pinned. The shells disagree on `exec -- cmd`: GNU bash 5.2.21
runs cmd, while dash tries to execute "--" (probe on this host). The
kernel currently takes `--` as the end of options and reads qmd.
Observed before the fix: "exec -- qmd mcp" => ["qmd/qmd"].

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Read any dash word after exec as an unresolved program

Follow U1 design 2.2: only the POSIX exec form is read, so a word that
starts with '-' after exec, `--` included, leaves the program
unresolved. bash 5.2.21 runs the command after `exec --` while dash
executes "--" itself, so neither reading is safe to assume. The failing
case from 1fcc9f23 now passes (node suite 107/107).

Also cover the inline npx server of an mcporter call
(openclaw/mcporter@93e0916c src/cli/ephemeral-target.ts:134-158): a
--server value that is an npx command line is an ad-hoc stdio server,
and another blank-holding name reads (other). These cases were added
after the implementation, so they are coverage, not failing-first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add failing-first Codex fixtures for shell call states

U1 design 6 fixtures (i)-(vii) run rollout records through
measure_codex_records, the Node bridge and the kernel's cli_lanes and
proxy. Record shapes follow openai/codex rust-v0.157.1 (36650394):
CommandExecutionItem and its snake_case status (protocol/src/items.rs
:199-285), exec end states (core/src/tools/events.rs:529-573: a
rejection is declined with exit -1) and the unified exec response
header (core/src/tools/context.rs:524-575). A rollout output never
carries the success flag (protocol/src/models.rs:2173-2182).

Before the adapter change 6 of 8 subtests fail:
- (iii) code-mode item declined: qmd succeeded 1, expected failed 1
  and not_executed 1; the same for a direct call's declined item;
- (iv) legacy exec_command with "Process exited with code 1":
  rtk_proxy succeeded 1, expected failed 1;
- (v) "Process running with session ID 3": qmd succeeded 1, expected
  unknown 1;
- (vi) no header and no item: qmd succeeded 1, expected unknown 1;
- (vii) local_shell_call running stack.json:280:
  codebase-memory-mcp succeeded 1 and downstream codebase-memory
  succeeded 1, expected unknown 1 in both.
(i) nested rtk proxy (proxy.calls 1, proxy.nested 1, succeeded 1) and
(ii) a failed item already pass. Must-stay controls pass: an exit 0
header, an exit line inside the output, a completed item that beats
a running header, a failed item with its header, and the aggregate
passing cli_lanes and proxy through with actors_with_success.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Map declined items and exec headers to Codex call states

U1 design 6: the Codex adapter now hands the kernel's callState what a
rollout shows about a shell call's outcome, so cli_lanes stops reading
declined or failed legacy commands as succeeded.

- A declined CommandExecution (openai/codex rust-v0.157.1 36650394
  core/src/tools/events.rs:562-573, exit -1) sets is_error and
  native_state "declined" on both result paths: the call's own
  function_call_output and the item's aggregated output. The kernel
  counts it failed and not_executed.
- A Bash-mapped call (exec_command, shell_command, shell or
  local_shell_call, chosen by call kind, never by output text) with no
  persisted item state is read by the unified exec header of
  core/src/tools/context.rs:524-548, and only by the lines before
  "Output:", in a linear line scan with a strict ASCII exit code.
  Exit 0 is success, another exit is_error. A running line (it wins
  over an exit line, unified_exec/process_manager.rs:1066-1071), a
  header with neither line or text without the header sets
  native_state "unknown": a rollout output never carries success
  (protocol/src/models.rs:2173-2182), and legacy history mode keeps no
  item (rollout/src/policy.rs:94-112). No shell, shell_command or
  local-shell handler exists at that revision, so those outputs read
  unknown unless an item state decides; apply_patch's "Exit code: N"
  text is not read.
- codex_call_name shares the MCP namespace rule between the call loop
  and the shell-call set, so an MCP tool named exec_command is not a
  shell call.

The failing-first fixtures from 4a7aa75a now pass (the three Python
suites: 160 tests OK). The header unit test
(test_exec_header_state_reads_only_the_pinned_header) was written
after the parser and is coverage, not failing-first. The lanes CLI
still writes nothing to stderr, and the Codex aggregate passes
cli_lanes and proxy through unchanged.
tools/skill-usage/README.md replaces the stage B sentence that said
these calls still read succeeded.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Regenerate workflow SHA256SUMS for the U1 kernel and tests

The N1 fix and the CLI-lane reading changed child-usage.mjs and
test-child-usage.mjs; stage C leaves the kernel byte-identical to
0c421c66 (sha256 acc7bb51...). Regenerated with the documented
command, `sha256sum -- *.mjs *.js *.json > SHA256SUMS` in
examples/claude-native/workflows, and `sha256sum --check --strict
SHA256SUMS` passes for all 13 files. Only those two digests change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Re-register the U1 files in manifests/evidence.json

The hot-file protocol (docs/lanes.md:94-103) puts the shared manifest
edit in the branch's last commit and re-registers every changed file
that the manifest already lists, with host_receipts.register_file.
scripts/validate.py checks SHA-256 and byte counts for each of them
and failed before this commit ("SHA-256 mismatch", "byte count
mismatch" for the workflows README, child-usage.mjs,
test-child-usage.mjs, tests/test_token_measurement.py and
tools/skill-usage/README.md). The eight re-registered entries are the
two READMEs, SHA256SUMS, child-usage.mjs, test-child-usage.mjs,
skill_usage.py, test_skill_usage.py and test_token_measurement.py.
Only their sha256 and bytes change, and validate.py now passes
(status passed, 7986 hashed files).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add failing-first tests for the U1 pivot loader and M4 repairs

Tests only, run against 7091a300 before any fix (U1 pivot brief D1, D6-D9,
D11; GPT-6 findings #8, #9, #11, #12, #13; Claude review R1, R2, R3, R9).

- D1 loader (test-child-usage.mjs, test_skill_usage.py): pin file contents,
  loadShellParser statuses (installed, not_installed, hash_mismatch),
  directory order, --shell-parser, commandInvocations null and cli_lanes
  parser_unavailable with proxy prefix_fallback, aggregate statuses,
  withShellTree freeing its tree, and the Node bridge awaiting the loader.
- D6: shell_script quotes argv with shlex.join, and -c is a script only for
  a shell; exec_header_state bounds the exit code to 9 digits (GPT-6 #11,
  #12 inputs verbatim).
- D7: outer-shell substitutions in a double-quoted run string are executed
  text (GPT-6 #8 verbatim), with single-quote, comment, backquote, sh, eval
  and ssh variants and must-stay controls that count once, not once per view.
- D8: n = 8000..64000 doubling for unclosed ((, $((, "$( ((, and ( << E
  runs, at most 2.5x per doubling and under 150 ms at 64000.
- D9: the stress check lets an exception fail it, includes an unquoted
  '$('.repeat(3000) input (NESTING_LIMIT), and a mutation control (a reader
  that throws for long input) shows the check now fails.
- R1, R3: a heredoc after a closed "$( )" keeps its reader; an escaped blank,
  semicolon or newline before a # is not a comment start.

Every expected value of the D7, R1 and R3 commands was run under GNU bash
5.2.21 and dash with a stub curl (and a stub ssh that runs sh on its stdin,
the remote login shell): both shells agreed on every count.

At 7091a300: node suite 115 passed, 31 failed (5 linear, 26 parser); D7 9
commands, R1 5, R3 3 fail with 0 != 1 on all three carriers; exec_header_state
raises ValueError on 5000 digits; shell_script flattens argv. The old D9
assertion passes when 7 of 7 stress inputs throw.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Quote Codex argv elements and bound the exec header exit code

U1 pivot brief D6 (GPT-6 findings #11 and #12).

shell_script joined an argv array with spaces, so a metacharacter inside one
element became shell syntax that was never executed: ["echo", "qmd; rtk proxy
qmd status"] counted one qmd call and one rtk proxy call. It now joins with
shlex.join (Python 3 shlex.join, the inverse of shlex.split), so each element
is one quoted word. The [shell, -lc, script] form (openai/codex rust-v0.157.1
codex-rs/core/src/shell.rs) stays the script itself, but only when the program
is a shell: -c of `echo -c ...` is one of echo's arguments.

exec_header_state passed the exit code text to int() whenever it was ASCII
digits, and CPython raises ValueError above 4,300 digits (int max str digits,
docs.python.org/3/library/stdtypes.html#int-max-str-digits), which stopped the
scan on one malformed header. The code is an i32 written in decimal (at most 10
digits with its sign, codex-rs/core/src/tools/context.rs:534-540); a code of
more than 9 digits is no native header and reads (False, "unknown").

The D6 tests of the previous commit pass: exec_header_state on 5000 digits,
the shell_script cases, and the three Codex bridge cases (local shell call,
code-mode item, `echo -c`). The existing Codex state tests still pass.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Scan unclosed (( in linear time: one paren table per text (D8, R9)

U1 pivot brief D8 (GPT-6 finding #9; Claude review R9). Every unclosed "((" or
"$((" made step() look ahead to the end of its line (indexOf the newline, then
count parentheses), so n = 8000, 16000 and 32000 of "((" took 246, 1218 and
4888 ms at 7091a300 (four times per doubling), and openers() filtered its whole
cut list once per heredoc operator.

parenClose(s) now matches every "(" of a text with its ")" in one stack pass per
line (bash(1) ARITHMETIC EVALUATION: (( )) and $(( )) count every ( and ) up to
the end of the line, which is classic parenthesis matching), and step() reads
its answer from the table. The last eight texts stay memoized so the scans that
alternate between a text and its slices do not rebuild it; executedTrace clears
the memo when a top-level analysis ends. openers() finds each operator's bounds
in the ascending cut list by binary search instead of filter and find.

Output is unchanged: executedText (both modes) and fetchKind of the base
kernel and this one agree on all 150,000 checks (the 243 covering commands, the
10,000 heredoc fuzz strings, and two seeded 20,000-string token corpora with
and without heredocs; 0 differences, 0 throws). Timings at n = 8000, 16000,
32000, 64000: "((" 3/6/13/27 ms, "$((" 4/9/19/40 ms, unclosed (( inside "$( "
5/8/18/38 ms, "(<<E" runs 8/14/30/64 ms, each within 2.5x per doubling and
under 150 ms at 64000. The D8 test allows 5 ms of timer noise on the 2.5x
bound (a linear scan showed 7.6 -> 19.4 ms once) and takes the best of five.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Read a # after an escaped blank or metacharacter as data (R3)

Claude review R3 of 3cb7c4f6. In a "$( )" frame step() treated a # as a comment
start whenever the previous character was a word break, ignoring whether that
break was escaped. `x="$(echo a\ #b)"; curl ...` therefore read `#b)"; curl ...`
as a comment: closeQuote and matchParen ran to the end of the command and the
`; curl` fell inside double-quoted data (shell_fetch 0, fetchKind null, one
unconfirmed mention). bash(1) COMMENTS: a word beginning with # starts a
comment, and an escaped blank, `;`, `(` ... or newline (a line continuation)
does not end a word (QUOTING: a backslash escapes the next character).

escapedAt(s, k) counts the backslashes before s[k]; an odd count means it is
escaped. It runs only for a # that follows a word break, and reads one run of
backslashes per such #, so the scan stays linear.

Verified with GNU bash 5.2.21 and dash and a stub curl: for each of the six
R3 inputs (escaped blank, escaped semicolon, backslash-newline, an even
backslash count before a real comment, a real comment inside "$( )", and a
top-level escaped blank) both shells ran the stub curl once, so each is one
executed fetch.

R3 test passes on all three shell carriers. Differential of the previous
commit's kernel and this one over 90,243 inputs (243 covering commands,
10,000 heredoc fuzz strings, and two seeded 40,000-string token corpora, with
and without heredocs; executedText in both modes and fetchKind): 228 differing
inputs, every one containing an escaped word break before a # (0 unexplained).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add failing-first R1 shapes and the --shell-parser refusal test

Tests only. Extends the R1 cases of 62cb6935 with the shapes the fix has to
cover beyond the reviewer's repro, each run under GNU bash 5.2.21 and dash with
a stub curl (both shells agreed):

- a quote inside the substitution: FOO="$(echo "a b")" bash <<'EOF' and the
  ssh variant, where the naive word split cuts the quoted word at the inner
  blank (curl ran once; with cat as the reader it ran 0 times);
- a string that runs over lines before the operator: x="a\nb" bash <<'EOF' and
  the single-quoted variant, and a "$( )" that closes on a later line with text
  after its ")" (curl ran once each);
- two closed "$( )" words, and controls that already read correctly at
  f83f8367 (a separator before the operator, a backquote span, a nested
  "$( )" with its own heredoc, cat as the reader).

At f83f8367 the 11 executed shapes each fail on all three shell carriers (33
subtests: shell_fetch 0 != 1), the four controls pass. Also: an explicit
--shell-parser that names a directory with no install exits 2 with the reason
and no report (the --rtk-db and --exceptions convention), never a silent
fallback to the default.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Bound each heredoc by the cuts of its own frame level (R1)

Claude review R1 of 3cb7c4f6. openers() kept one cut list for every nesting
level, so the ( and ) that a double-quoted "$( )" adds also bounded the heredoc
operators outside it: in FOO="$(pwd)" bash <<'EOF' the operator's command
started at the closing quote, its reader resolved to the empty string, and a
body that bash reads as a script became data (shell_fetch 0, one unconfirmed
mention, fetchKind null; real bash and dash run the stub curl once).

step() now reports each cut with its frame (on.cut(i, frame): the "$(" frame
that a ( opens and a ) closes, else the frame the unit sits in, undefined for
the top level) and openers() keeps one ascending cut list per level; an
operator is bounded by the cuts of the level on top of the stack, found by
binary search. The cuts of a substitution bound the commands inside it and no
longer the command that contains it. commandWords() reads a "$( )" or a
backquoted span inside double quotes as one unit of its word (POSIX.1-2024 XCU
2.6.3: the substitution's text is tokenized on its own), so a quote or blank
inside it does not split the word: FOO="$(echo "a b")" bash <<'EOF' names bash.

R1 tests (62cb6935, 58502c60) pass on all three shell carriers, and the lane
reading of FOO="$(pwd)" bash <<'EOF' with qmd in the body counts one qmd call.
The shapes below are documented limits, pinned by their own test: a command that
a quote or "$( )" carries over lines (x="$(cat <<'A' ... \n)" bash <<'B'
and x="a\nb" bash <<'EOF') has its head on an earlier line, so this line alone
gives the operator no reader and the body reads as data (fetch_mentions_unconfirmed
1, status incomplete: a possible fetch, never a lost one). A first attempt cut
the line at the closing quote's word end; measured on the fuzz corpus it named
the wrong owner as often as the right one (ssh host cat "a\nb" bash <<EOF), so
it was dropped.

Differential of f83f8367 and this commit over 90,243 inputs (243 covering
commands, 10,000 heredoc fuzz strings, two seeded 40,000-string token corpora,
without and with heredocs; executedText in both modes, fetchKind and the M4
confirmed and unconfirmed counts): 17 inputs differ, all with a quote,
backquote or "$( )" before a heredoc operator (0 unexplained); no input lost
mass (confirmed plus unconfirmed); 16 keep both counts, one moves a fetch from
confirmed to unconfirmed. That one, bash -s -- "$(pwd)"bash <<EOF with an
unquoted-delimiter body, now names bash as the reader (correct) and reads the
body as source, where a substitution inside single quotes is data: the
"$( )"-free twin reads C=0 U=2 at 7091a300 as well (real bash runs curl once:
the outer shell expands an unquoted heredoc body before the reader sees it, a
limit that predates R1 and is unchanged).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Read a double-quoted run string in two views: outer expansion (D7)

U1 pivot D7, GPT-6 finding #8. N1 read a double-quoted string that a shell runs
only as that shell's input, so a substitution the OUTER shell runs while it
expands the word was lost whenever the inner shell reads it as data: a quoted
heredoc, single quotes, a comment. bash -c "cat <<'EOF'\n$(curl https://example.org)\nEOF"
went from fetchKind 'fetch' (cf3fb72e) to null with shell_fetch 0. bash 5.2.21 and
dash run the stub curl once: the expanding shell runs each unescaped "$( )" and
backquote before the shell it starts reads anything (POSIX.1-2024 XCU 2.2.3
Double-Quotes, 2.6.3 Command Substitution). Two views now: outer bodies first,
each as commands of their own, then the inner shell's read of the string with
each outer span replaced by a placeholder word (its output is unknown), so a
substitution both views could see counts once. A backslash-escaped \$( ) or
backquote is no outer span and stays the inner shell's (bash -c "echo \$(curl u)"
runs curl once, in the inner shell). Single-quoted run strings have no outer
expansion and are unchanged.

outerSpans() is the one scanner of a double-quoted string's substitutions (escape
pairs skipped, "$((" arithmetic not a span but the substitutions inside found,
NESTING_LIMIT and unterminated spans as before); quotedData() now consumes it, so
the data view and the run view cannot disagree. The ssh remote ranges (marks)
cover the literal text between spans, not the local substitutions.

D7 tests (62cb6935) pass on all three carriers: 13 executed, 5 inner and 2
twice-run commands, 2 data commands; the qmd lane test counts one call.

Evidence, all against GNU bash 5.2.21 and dash with stub curl/wget/ssh (ssh runs
its arguments with sh -c, or reads stdin as the remote login shell), counting
stub calls against the kernel's confirmed count (shell_fetch + loopback):
- compositional oracle, 360 commands (4 double-quoted run carriers x 10 contexts
  x 4 atoms, two-atom strings, single-quoted carriers, heredoc readers x
  delimiters x bodies): base 7091a300 301 exact, 59 under, 0 over; 43b9d947 329,
  31, 0; this commit 353, 7, 0. The 7 left are one class, a shell-read heredoc
  with an unquoted delimiter whose body has a substitution inside single quotes
  (the outer shell expands the body first), next commit.
- random valid-bash fuzz, 40,000 candidates, 7,883 accepted by bash -n and dash -n
  with equal stub-call counts under both (361 with a fetch): base, f83f8367,
  43b9d947 and this commit agree on the same 7,789 (98.81%) and the same 94
  disagreements (12 over: word glue and a trailing backslash; 82 under: remote
  ssh commands and glued words), all present at 7091a300.
- differential of 43b9d947 and this commit over the covering commands and three
  fuzz corpora (90,243 inputs): 13 inputs lose confirmed mass, every one invalid
  bash in at least one shell; of the 24 differing inputs without HTTP-library or
  gh tokens, 21 are invalid, 2 valid inputs move closer to real bash and 1
  (xcat &&bash -c "'$(curl u)'") counts a substitution that a failing && never
  reaches, which no static reading models.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add failing-first tests: an unquoted heredoc body is expanded first

Tests only. D7 for a heredoc body (U1 pivot D7, same defect class as GPT-6 #8):
with an unquoted delimiter the shell that reads the command line expands the
body before the shell that reads the heredoc sees it, so each unescaped
substitution runs in the outer shell whatever its quotes or a comment say
(bash(1) Here Documents; POSIX.1-2024 XCU 2.7.4), quotes are literal, and a
backslash escapes only $ ` \ and a newline. 20 commands, each run under GNU bash
5.2.21 and dash with stub curl and ssh: curl ran once for 17 (a single-quoted
substitution, a comment, nested quotes, a substitution over lines, a <<- body
with tab indents, arithmetic, an ssh reader, a run string in the body, and
controls that already read correctly), twice for one (an outer and an inner
substitution), and never for two (a quoted delimiter passes the body as is).

At 67a4acbb 9 of the 17 fail on all three shell carriers (shell_fetch 0 != 1,
one unconfirmed mention: bash <<EOF with echo '$(curl u)', the FOO="$(pwd)" and
ssh readers, a comment, nested quotes, a multi-line substitution, <<-, and
arithmetic); the rest and the data commands pass.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Expand an unquoted heredoc body a shell reads before reading it

The heredoc form of D7 (GPT-6 #8: the outer shell runs what it expands before
the inner shell sees the text). With an unquoted delimiter the shell that reads
the command line expands the body first: each unescaped $( ) and backquote runs
there whatever its quotes or a comment say, quotes are literal in a body, and a
backslash escapes only $ ` \ and a newline (bash(1) Here Documents; POSIX.1-2024
XCU 2.7.4). resolveHeredocs() analyzed a body that a shell (or in inlineHttp
mode an interpreter) reads only as that reader's source, so bash <<EOF with
echo '$(curl u)' read as data although real bash and dash run curl once, and
so did a substitution in a comment, nested quotes inside it, one over lines,
arithmetic, and a <<- body: the same 9 shapes as the failing-first commit
b6f3b454. R1 made the FOO="$(pwd)" bash <<EOF form read this way too, so this
also removes the accident that had hidden the limit there.

The body now has two views, sharing D7's outerSpans/withoutSpans/outerBody: the
outer spans first, as commands of their own (heredocs inside them resolved
there), then the reader's source with each span replaced by the placeholder word
and its own escapes applied (heredocText). A quoted delimiter is unchanged
(nothing is expanded), and so is a data heredoc such as cat <<EOF, which keeps
the substitutions its regular expression finds. The ssh remote ranges cover the
body's literal text, not its local substitutions.

The heredoc test of b6f3b454 passes on all three carriers (17 once, 1 twice, 2
data commands); tests.test_token_measurement passes (59). Evidence, real bash
5.2.21 and dash with stub curl/wget/ssh, kernel confirmed count against stub calls:
- compositional oracle, 441 commands (D7's 360 plus escaped, comment and
  nested-quote bodies): base 7091a300 368 exact, 73 under; 67a4acbb 420, 21;
  this commit 441 of 441 exact, 0 under, 0 over.
- grammar fuzz, 6,000 valid nested commands (statements, sequences, substitutions,
  backquotes, single and double quotes, run strings, heredocs with unique
  delimiters) that bash and dash ran with equal stub-call counts (3,049 with a
  fetch, 1,597 with a run string, 1,471 with a heredoc): base 5,788 exact
  (96.47%), 212 under, 0 over; 67a4acbb 5,970 (99.50%), 30 under, 0 over; this
  commit 5,975 (99.58%), 25 under, 0 over. All 25 are one older gap, a run string
  right after a backquote (echo `bash -c "curl u"`), whose text is invisible to
  the raw detector as well; fixed next.
- random valid-bash fuzz, two seeds of 40,000 candidates (7,883 and 7,838 valid,
  361 and 374 with a fetch): base, 67a4acbb and this commit agree on every
  input.
- differential of 67a4acbb and this commit over 90,243 inputs: 924 differ, 20 with
  a lower confirmed count, of which 19 are invalid bash or valid only with a
  heredoc-at-end-of-file warning, and the one clean input (a backquote span that
  holds bash <<EOF) moves from 3 to 0 confirmed where real bash and dash run
  curl 0 times.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add failing-first tests: a run string after an opening backquote

Tests only. RUN_QUOTED anchors a shell, eval or ssh word at the text start, a
blank, ; & | or (, but not at a backquote, so bash -c "curl u" inside `...` read
as double-quoted data. The fetch was lost, and the raw detector, which anchors
curl itself and not the quote, counted no possible fetch either (C=0 U=0). Found
by the grammar fuzz of 2a33ef90: all 25 commands left under-counted there were
this one shape.

8 commands run under GNU bash 5.2.21 and dash with stub curl and wget: the stub
ran once for each (bash -c and sh -c with double or single quotes, eval, an ssh
string, a backquote inside a single-quoted run string, an escaped backquote in a
double-quoted one, and two controls that already read correctly: a blank after the
backquote and $( ) in its place) and never for echo `echo "curl u"` (data).

At 2a33ef90 the 6 backquote-adjacent shapes fail on all three shell carriers
(shell_fetch 0 != 1, no unconfirmed mention: 19 failures with fetchKind).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Anchor a run string at an opening backquote too

RUN_QUOTED and RUN_HTTP_CODE anchored a shell, eval, ssh or interpreter word at the
text start, a blank, ; & | or (, but not at the backquote that opens a `...`
substitution, so bash -c "curl u" inside backquotes read as double-quoted data: the
fetch was lost and no possible fetch was counted either (C=0 U=0). bash(1)
Command Substitution: a `...` body is shell text with its own command positions,
as a $( ) body is; the anchor class now holds the backquote (one character in
each regular expression; no new alternative, so no new backtracking).

The 6 backquote shapes of the failing-first commit (66550644) pass on all three
carriers; tests.test_token_measurement passes (60).

Evidence, real bash 5.2.21 and dash with stub curl/wget/ssh, kernel confirmed count
against stub calls: the 441-command oracle stays 441 exact; the grammar fuzz that
found the gap is now exact on every input, 6,000 of 6,000 (3,049 with a fetch,
1,597 with a run string, 1,471 with a heredoc; 2a33ef90 5,975 exact, 7091a300
5,788, 96.47%) and, on a second seed, 11,999 of 11,999 (base 11,608, 96.74%, 391
under; 0 over throughout). Differential of 2a33ef90 and this commit over 90,243
inputs: 216 differ, 1 with a lower confirmed count, an unterminated string (invalid
bash).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add a pinned, verified tree-sitter-bash loader that fails closed (D1)

U1 pivot D1. The CLI-lane reading is to run on tree-sitter-bash, the grammar
OpenAI Codex uses for the same job (openai/codex rust-v0.157.1 36650394
codex-rs/shell-command/src/bash.rs; tree-sitter-bash = "0.25" at
codex-rs/Cargo.toml:522, both read at the tag). This commit adds the seam and the
gate; the AST reading itself is the next stage.

shell-parser.pin.json pins the install: web-tree-sitter 0.27.0 and
tree-sitter-bash 0.25.1 with their npm integrity values (checked against the
registry's dist.integrity), upstream tags and commits (v0.27.0 6070dbfe =
npm gitHead; v0.25.1 names a06c2e44 while the package was published from
80132668, five minutes earlier, recorded as a note), the sha256 of the six
files the loader reads, and the install command (npm install --prefix <dir>
--ignore-scripts --no-audit --no-fund --save-exact). loadShellParser(dir)
resolves the directory as the argument (the new --shell-parser flag), then
CHILD_USAGE_SHELL_PARSER, then the ecosystem tools directory under HOME, reads
each pinned file once, compares its sha256 and the lockfile's version and
integrity values with the pin, and executes only those bytes: the runtime module
is imported from a private 0600 copy of the hashed bytes and the two wasm files
are handed over as bytes, so the code that was hashed is the code that runs.
Results carry no path: { ok, versions, wasm_sha256 } or { ok: false, reason }
with not_installed (directory, a pinned file or the lockfile missing),
hash_mismatch, load_error (an unreadable file, the pin file, init) and
not_loaded (this process never awaited the loader, so a forgotten await is not
mistaken for a missing install). A failed load leaves no parser, whatever an
earlier call did; a successful one reuses the initialized runtime. withShellTree()
frees every tree in a finally.

Fail closed: commandInvocations() returns null without a parser and measurement
reports cli_lanes { status: parser_unavailable, reason }, keeps the prefix rule for
the rtk proxy carrier (proxy.rule prefix_fallback, invocations and in_ctx_code
null) and never falls back to the text scanners for lane counting; M4 stays text
based. A measured cli_lanes carries status measured and the parser record (versions
and the two wasm sha256 values); aggregates report measured, incomplete (mixed
actors, with the counts), parser_unavailable or not_measured, and sum proxy
counters null-aware. The CLI awaits the loader; a --shell-parser that cannot be
honored exits 2 with the reason (like --rtk-db), any other missing install only
leaves cli_lanes unavailable. The Codex bridge awaits the loader too.

Interim, until the AST reading lands: with a verified parser loaded,
commandInvocations() still runs the scanner-based lane layer of 3cb7c4f6, so every
existing lane fixture reads as before; the gate, the status fields and the seam
(withShellTree) are final.

Tests (62cb6935, 58502c60 first, run failing at 7091a300): the pin's contents;
not_installed for an absent directory, a missing pinned file and a missing
lockfile; hash_mismatch for one flipped byte in the bash wasm, a modified
web-tree-sitter.js that is never imported, a modified package.json, another
integrity value and another version in the lockfile; load_error for an unreadable
file; directory order; --shell-parser; a good install loading again after
failures; unavailable and mixed aggregates; the bridge in both states. The lane
fixtures load the parser first and skip with a message on a host without an install.
tests: node suite 149 passed; tests.test_token_measurement, test_skill_usage and
test_child_usage_suite 175 passed. D9's stress check reports the TypeError the
lane checks threw before the parser was loaded (threw: TypeError x8), which the
old assertion would have passed.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add failing-first timing tests for quote-heavy scripts (D8)

Tests only. D8 asks the M4 text scanners for linear time; the unclosed "((" input
of GPT-6 #9 is fixed (5e814331), but profiling a realistic 346 KB script (4,000
lines of echo "step N: $(date)" >> log; qmd search "term N" | head) at e5238f39
showed the same quadratic shape from another term: 51% of the time in the
run-string regular expression and 33% in scanQuotes, both from one read of the
whole text built so far per quote (the regex, and the string flattening that
indexing a concatenation rope forces). A 42 KB script takes 47 ms, 84 KB 108 ms,
172 KB 467 ms, 346 KB 2,078 ms and 694 KB 8,236 ms a scan, and the kernel scans
each shell call several times.

Five doubling checks (n = 8000..64000, best of five, at most 2.5x per doubling
plus 5 ms of noise, 64,000 under 1.5 s): a run of quoted words, a run of
double-quoted substitutions in inlineHttp mode, a run of `bash -c "x"` strings,
a run of words with a # inside, and a long script of echo, substitution and pipe
lines through fetchKind. At e5238f39: 61, 230, 902, 5,861 ms; 168, 651, 3,728 ms;
991, 4,486 ms; 9, 32, 120, 459 ms; 89, 392, 1,930 ms (the series end where a run
passes 1.5 s). All five fail.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Read only the tail of the built text at each quote (D8: linear scripts)

scanQuotes() tested RUN_QUOTED (and in inlineHttp mode RUN_HTTP_CODE) against the
whole text built so far at every quote. Reading out.s for that flattens the string
concatenation and scans the full prefix each time, so a script with n quoted
words cost n squared. CPU profile of a 346 KB script (4,000 lines of echo "step N:
$(date)" >> log; qmd search "term N" | head) at e5238f39: 51% in the run-string
regular expression, 33% in scanQuotes itself (the flattening), 10% in GC. The
five timing tests of 9da57414 failed on it.

scanQuotes now keeps the end of its output in `tail`, cut back to 512 characters
whenever it passes 1,024, updated at every write (put and putTrace; quotedData
builds into its own buffer first), and reads only that: the # test uses the last
character of the tail, the emptiness test out.p.length. Once a cut has been made
the two detectors run without their start-of-text alternative (a cut is no start
of text). Limit, stated in the code: a run keyword more than 512 characters
before its quote, after a text that has grown past 1,024, is missed and the
string reads as data. Only a single ssh option word that long can put it there
(ssh -oProxyCommand=<1500 characters> host "curl u": fetch before, null after;
100, 400, 520 and 700 characters read the same), since the words between a
shell, eval or ssh and its string are short options and one host.

Timings (best of five, ms), n = 8000, 16000, 32000, 64000: quoted words 61, 230,
902, 5,861 -> 8, 15, 31, 69; double-quoted substitutions (inlineHttp) 168, 651,
3,728 -> 15, 31, 67, 140; run strings 991, 4,486 -> 38, 83, 173, 358; words with
# 9, 32, 120, 459 -> 3, 6, 14, 37; a long echo/substitution/pipe script through
fetchKind 89, 392, 1,930 -> 18, 34, 71, 154. Realistic scripts, executedText: 84
KB 108 -> 37 ms, 346 KB 2,078 -> 100 ms, 694 KB 8,236 -> 205 ms, 1.4 MB 393 ms
(2x per doubling); the node suite runs in 12 s instead of 41 s.

Output is unchanged: executedText (both modes) and fetchKind of e5238f39 and this
commit agree on all 1,084 shell texts of this repository (822 .sh files and fenced
shell blocks of .md files, 1.66 MB, largest 82 KB, 83 over 4,000 characters), on
the 243 covering commands and the three fuzz corpora (90,000 inputs), and on 600
synthetic long scripts (0 differences, 0 throws); the 441-command oracle stays 441
exact, and the 6,000-command grammar fuzz stays 6,000 exact (real bash 5.2.21 and
dash with stub curl).

Also fixed while rewriting the call: the ssh remote offset test listed the
separators before RUN_QUOTED's keyword but not the backquote that 296c4f1d added,
so in echo `ssh host "qmd get a"` the qmd call read as local; it now reads
qmd/qmd remote (fixture added; the same input read local at e5238f39).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Flip one real byte in each pinned file in the loader tests

The loader tests changed the bash grammar wasm by reading it as UTF-8 and
appending a NUL, which rewrites the whole binary: the test passed because the
file differed, not because one byte did. It now reads a Buffer and flips bit 0 of
the middle byte in place (same length, so only the hash can tell), and does the
same for each of the six pinned files in turn: every one gives hash_mismatch,
leaves no parser (commandInvocations returns null, the scanners do not step in) and
runs no unverified code (the modified web-tree-sitter.js sets a global if it is
ever executed; it is not). The missing-file and missing-lockfile cases now also
assert that no parser stays loaded. Node suite: 160 passed (149 before).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add a failing-first test: a run keyword past the tail window

Tests only. The brief's D8 says a work budget "marks the command unclassifiable for
M4 instead of throwing". 1d77c34c bounds scanQuotes to the last 512 characters of
its output, but an ssh option word long enough to push `ssh` out of the window
(in a text past 1,024 characters) drops the run string silently: the fetch in
ssh -oProxyCommand=<1500 characters> host "curl u" is read as data with no
possible fetch counted either (shell_fetch 0, unclassifiable 0, unconfirmed 0),
which a budget must not do.

The test: option words of 100, 400 and 700 characters read as before (the fetch
confirmed, unclassifiable 0); 1500 characters counts one unclassifiable operation
(remote_fetches 1, no confirmed shell fetch); the count is once per command and
only for the command that held the keyword (two later quoted commands add nothing);
a long printf with hundreds of words and quotes but no run keyword adds none.

At 1d77c34c the two 1500-character cases fail on all three shell carriers
(unclassifiable 0 != 1: 6 failures); the 100, 400 and 700 cases and the printf
control pass.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Count a run string the tail window lost track of as unclassifiable

The brief's D8 fallback for input past a work budget is "marks the command
unclassifiable for M4", not a silent miss. 1d77c34c bounds scanQuotes to the last
512 characters of its output; an ssh option word long enough to push `ssh` out of
the window (in a text past 1,024 characters) dropped the run string with no trace:
ssh -oProxyCommand=<1500 characters> host "curl u" read as shell_fetch 0,
unclassifiable 0, unconfirmed 0.

scanQuotes now notes, outside quotes, when a shell, eval, ssh or su word (and its
blank) has been written since the last command separator (a last-letter pre-check
keeps it off the hot path). When a cut has left the tail without that word and
without a separator, the next quote that does not read as a run string is counted
once in windowMisses, which countFetches adds to unclassifiable (an operation of
unknown kind: remote_fetches 1, no confirmed shell fetch). The keyword is noted at
write time because the discarded text cannot say whether a word was quoted: a first
version that looked for the word in the discarded text also fired on 4 real scripts
of this repository (adoption/bootstrap-linux.sh, bootstrap-macos.sh,
launchd-agents.sh and a memory-stack README block), where separators inside quoted
data are blanked and a word such as sh sat in a long quoted list.

The test of 0b2b5778 passes on all three carriers (100, 400 and 700 characters
read as before; 1500 counts one unclassifiable, once for the command, and adds
nothing for later commands or for a long printf with no run keyword).

The counter is zero on every corpus: complete m4 results of the pre-window kernel
(e5238f39) and this one are identical on the 243 covering commands, 10,000 heredoc
fuzz strings, two 40,000-string token corpora and all 1,084 shell texts of this
repository (822 .sh files and fenced shell blocks, 1.66 MB), 2 tools each; the 600
synthetic long scripts agree; the 441-command oracle is 441 exact and the grammar
fuzz 11,999 of 11,999 exact against real bash and dash. Realistic scripts, executedText:
84 KB 50 ms, 346 KB 125 ms, 694 KB 285 ms (2,078 and 8,236 ms at e5238f39).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Keep 150 ms for the brief input only; loosen the extra timing shapes

The brief's D8 requires at most 2.5x per doubling and under 150 ms at n = 64,000 for
'(('.repeat(n) + 'qmd'. The timing tests added alongside it (a run of $((, the same
"$( ((, a run of ( << E, and the quote-heavy shapes) had inherited that absolute
bound. Three of them run 64-80 ms at 64,000 against 150 ms on this host, thin
headroom for a slower CI runner. Only the brief's input keeps 150 ms (it runs about
30 ms); every other shape keeps the 2.5x ratio (plus 5 ms of timer noise, best of
five) and takes an absolute bound of 1.5 s at 64,000, which is still far below the
quadratic behavior they replaced (the same shapes take 1.9 to 10 s and fail the ratio
at 7091a300).

Run against the base kernel (7091a300) with this test file, all ten timing shapes
fail (for example (( 313, 1,219, 5,109 ms; quoted words 62, 239, 928, 6,211 ms), and
against this head all ten pass (for example (( 3, 7, 17, 32 ms; quoted words
8, 15, 32, 71 ms; run strings 42, 95, 205, 416 ms).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Retry the doubling timing rounds so one pause cannot fail a linear scan

A direct run of test-child-usage.mjs at da2d578d failed one assertion, "a run of unclosed
(( at n = 8000, 16000, 32000, 64000 takes [4, 7, 27, 31] ms": the 32,000 reading was 27 ms
where the scan takes about 15, so the ratio to the 7 ms before it (3.9x) broke the 2.5x
(+5 ms) bound. The same suite had passed inside the Python run seconds earlier. A flaky
timing check is a CI hazard, so the doubling measurement now takes the best of five per
size and, when the ratio still fails, up to two more rounds and the elementwise minimum of
all of them: a pause cannot fail a linear scan, while a quadratic one fails every round.

Checks on the final file: six consecutive runs and two runs beside six busy CPU loops, all
160 passed; against the base kernel 7091a300 all ten timing shapes still fail (the retry
does not hide quadratic behavior).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add failing-first tests for the AST command-position layer (D10, D11)

The GPT-6 review of the scanner-based lane reading (3cb7c4f6) found ten
command-position defects (#1-#7, #10 here; #8, #9, #11-#13 were repaired
earlier) and the Claude review five more (R4-R7). This commit holds their
probes before any repair, so each one fails for its defect.

- tests/test_token_measurement.py: one test per finding, with the GPT-6
  inputs verbatim, asserting what the finding states (lane calls, proxy
  calls, exclusions, unresolved programs, downstream keys, call states),
  not a display string. Also parse_errors (D2), the closed name vocabulary
  (D5) and mcporter's help/version tokens (R6: checked against
  openclaw/mcporter@93e0916c src/cli.ts:137-145 and flag-utils.ts:34-41,
  the code already matches; the test pins it).
- test-child-usage.mjs: a program name is emitted only for a lane
  executable or a name the reading interprets (GPT-6 #10).
- tests/test_command_position_oracle.py (D10): a deterministic generator
  and a hand-written probe list run under REAL bash 5.2 with env -i, a
  PATH of logging stubs, a temporary HOME and cwd and a 5 s limit; the
  multiset of lanes that ran must equal the lanes commandInvocations reads.
  Against the scanner layer: 38 of 99 probes and 59 of 457 generated
  commands disagree with the run.

Every expected value is the run of real bash and dash under stub
executables (the oracle covers each input), not what this kernel says.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Read shell text with the tree-sitter-bash AST for CLI lanes (D2-D4)

Replaces the scanner-based command-position layer (simpleCommands,
resolveInvocation, analyzeScript) with a walk of the tree-sitter-bash
tree, the grammar openai/codex rust-v0.157.1 codex-rs/shell-command
src/bash.rs uses for the same job. The option tables, runner maps and
mcporter grammar are unchanged and now operate on word values.

- D2: every `command` node counts wherever it sits (substitutions inside
  strings, arithmetic, arrays and unquoted heredoc bodies, subshells,
  compound and control statements, function bodies); array elements, case
  patterns, [[ ]] operands, function names and the words of any non-lane
  program never do. ERROR nodes and what is under them are skipped and the
  call counts once in cli_lanes.parse_errors.
- D3: wordValue() returns a literal or null: unquoted escapes, '...', the
  double-quote rule (a backslash goes only before $ ` " \ and newline),
  $'...' decoded as bash does, concatenations; any expansion, glob, brace
  expansion is unknown. A tilde prefix keeps the word known for program
  identity (only the directory changes: ~30 lane invocations by ~/path in
  136,361 real commands) but not as a script.
- D4: bash|sh|dash|zsh|ksh -c reads the first operand after the options
  (-n, -D and -o noexec run nothing; `--`, `-` end the options); a shell
  with no -c reads standard input, so a heredoc or here-string attached to
  the last command of a list, pipeline or negation, through env, timeout,
  nice, nohup, stdbuf, rtk proxy (not xargs, which consumes it), has its
  body read; eval and ssh join their words as the callee does; rtk proxy
  resolves its program with exec semantics (command and exec run nothing,
  time is unresolved, an expansion in its one argument is unresolved).
- Backquotes and outer substitutions of an unquoted heredoc body and of a
  double-quoted script are read from the raw text with the kept M4 helpers
  (tree-sitter-bash leaves backquotes in heredoc_content and mis-reads
  some bodies); a script with an expansion is read with a placeholder.
- The grammar folds the next line into a command in some shapes (an
  unquoted == before ;, &&, | or a newline; a trailing blank before a
  newline; `<<'EOF' 2>&1 | tail`): 732 of 136,361 real commands with no
  error and 163 with one. A separator ERROR or an unescaped newline inside
  one command's words starts a new command.
- simple counts assignment-only statements, declarations and each loop,
  if and case as well as commands (POSIX.1-2024 XCU 2.9.1, 2.9.4), so the
  ambiguity of existing fixtures is unchanged.

Fixtures: EMPTY_CLI gains parse_errors (D2), and `ssh host bash -s <<EOF`
now also reads the remote bash as a command of its own (ssh(1): the words
after the destination run remotely). Every other existing lane fixture
passes unchanged.

Against real bash 5.2: the oracle agrees on 457 of 457 generated commands
(581 lane runs) and 99 of 99 probes. Over 136,361 distinct real commands
the old and new readings agree on 136,167; the differences are old false
positives, unknown wrappers (setsid, bwrap) and grammar defects. Still
failing here by design, fixed next: D5 vocabulary (5 tests) and R7 states
(7 tests); the node suite fails only on the D5 program-name check.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Re-expect the closed name vocabulary in the lane fixtures (D5)

U1 pivot D5 and GPT-6 #10: commandInvocations returns `program` only for
a lane executable or a name this reading interprets itself (a wrapper, a
shell, eval, ssh, rtk), and an mcporter downstream key is emitted only for
the stack's own servers; every other name-shaped string (a host, an id, a
package, `linear`) reads null or (other). The old expectations that showed
`-/echo`, `-/git`, `-/pytest` or `@linear` are wrong under D5, so they are
changed here, before the code: 7 node checks and 3 Python tests fail
against the walker of 1d35df12 with the observed values recorded in the
log (`program` still holds "my-private-host.example", the downstream key
still holds "linear" and "call_PRIVATE").

The selector-form coverage the `linear` cases gave is kept with servers of
the stack (`socraticode`, `serena`, `jcodemunch`, `ai-memory`, `qmd`,
`headroom`, `context-mode`, `codebase-memory`), read by name in every form.
Sources of the closed set: manifests/stack.json:280 (codebase-memory),
:407 (context-mode), :629 and :966 (socraticode); the rest is the
coordinator's list in the pivot brief.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Emit only fixed names: a closed program set and mcporter server set (D5)

GPT-6 #10: `program` carried any name-shaped basename (a host, an id, a
package) and an mcporter downstream key any name that passed a character
class, so the no-id, no-host output contract held only by luck.

- `program` is set only for a lane executable or a name the reading
  interprets itself (the wrappers, the shells, eval, ssh, rtk); every other
  program reads null.
- An mcporter downstream key is one of codebase-memory, context-mode,
  jcodemunch, serena, socraticode, qmd, headroom or ai-memory (the
  stack's servers: manifests/stack.json:280, :407, :629, :966 and the
  coordinator's list), else (other), (http), (stdio) or (unresolved).

The tests of the previous commit now pass (node 162 of 162; the Python
name-vocabulary test and the three re-expected fixtures).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Expect an interrupted counter in lane and downstream rows (R7)

Claude review R7: a Claude Code result whose toolUseResult.interrupted is
true read as succeeded. The command started and was cut short, so it is a
state of its own. The shared fixture helpers (lane_row, downstream_row and
the Codex cli_lane_row and downstream row) now carry `interrupted: 0`; the
lane and downstream rows of the kernel do not have the key yet, so the
tests that compare whole rows fail (see the log of this run) until the next
commit adds the counter and the state.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Read not-executed calls where the client records them; add interrupted (R7)

Claude review R7: the marker "The user doesn't want to proceed" was matched
against the row's toolUseResult, where it was never observed; it opens the
result CONTENT. Observed on this host's Claude Code transcripts (122,648
Bash results in 4,355 files, count-only): the rejection content (25 rows,
toolUseResult "User rejected tool use"), a row-level toolDenialKind on
every client denial (permission-rule 750, user-rejected 46, cancelled 1,
automode-unavailable 1), the classifier text "The server-side auto mode
classifier gave no verdict" (its own text says the check failed, not the
action), and 209 rows of a host hook's refusal text ("This agent is
isolated") that carry no toolDenialKind and stay failed (open: its meaning
is the hook's, not the client's).

- callState: not_executed when the row has a toolDenialKind, or the content
  opens with <tool_use_error>, "The user doesn't want to proceed" or the
  classifier text, or toolUseResult names a hook denial, a permission
  denial or a rejection.
- toolUseResult.interrupted === true is its own state, `interrupted`
  (none of 25,506 observed object results, so it is a synthetic fixture),
  with a counter in lane and downstream rows; it no longer reads succeeded.

The R7 test of the first commit and the fixtures of the previous commit
pass (measurement and Codex modules: 188 tests).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Remove the lane-only marks plumbing from the M4 text machinery

The scanner-based lane reading asked executedTrace for `marks` (the raw
range of each ssh body and the offset of each command substitution of a
data heredoc body). The AST walker reads both from the tree and the raw
text, so nothing passes marks any more: the parameter and every branch
that wrote to it (resolveHeredocs, scanQuotes, outerBody, quotedData) and
rawRange are removed. The M4 readings are unchanged: the complete m4
object of measureTranscript (Bash and ctx shell carriers), executedText in
both modes and fetchKind are identical between the kernel before the
walker (f1ed98ac) and this one over 136,361 distinct real commands, the
243 covering commands, and 2 x 40,000 fuzz inputs (0 differences, 0
exceptions).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Test the recovery of commands the grammar joined to the previous one

The walker of 1d35df12 starts a new command at a separator ERROR or an
unescaped newline inside one command's words, because tree-sitter-bash
0.25.1 reads what follows a command as its arguments in these shapes
(measured on this host's transcripts: 732 of 136,361 distinct real
commands with no error and 163 with one): an unquoted == or =~ before ;,
&&, | or a newline, a redirection after a here-document operator followed
by | or ;, and a multi-line shape found in the real corpus. This commit
adds the tests that pin it:

- tests/test_command_position_oracle.py: 12 recovery probes whose expected
  lanes are what REAL bash 5.2 ran (12 of 12 agree; 10 carry a parse
  error, 2 none).
- tests/test_token_measurement.py: the same shapes through
  measureTranscript, with parse_errors counted once for the calls whose
  tree has an ERROR, and two controls that are not boundaries (a
  backslash-newline continuation and a comment).

Mutation control (recovery disabled in a scratch copy of the kernel): all
12 oracle probes fail (real bash ran the lane, the kernel read none or
one fewer, e.g. `echo ==; qmd get a` ran qmd, read []) and 9 of the 12
measurement cases fail with the observed values ({} != {'qmd': 1}).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add failing-first tests: a line that begins with a backslash

Found by the 14,461-command run of the oracle generator over more seeds
against real bash: a line that begins with a backslash (`\ls`, the alias
bypass; `\markitdown a.json`) reaches tree-sitter-bash 0.25.1 as a word that
begins with the newline, so the line joins the command before it, and when
it is the first line of a here-document body the grammar leaves the body
node without it and adds its words to the operator's arguments (which hid
that the owner is a shell). A here-string that precedes such a line was fed
to the last command of the joined node instead of its own.

Five shapes are added to the oracle's recovery probes and to the
measurement recovery test; real bash ran every lane in them. Against the
kernel of 907ac35f the two tests fail with the observed values (recorded in
the log of this run), e.g. `{ timeout 5 repomix` + newline +
`\markitdown --flag; }` read {'repomix': 1} where bash ran both, and the
shell heredoc whose first body line starts with a backslash read {}.

Also hardens the oracle generator: it drops a text bash rejects
(`bash -n`), keeps a function definition and its call in one group,
does not put `!` or a heredoc on the right of a pipe (`x | ! y` is not
POSIX; a reader that ignores its pipe kills the writer with SIGPIPE), and
writes data heredocs to /dev/null.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Read a line that begins with a backslash as a command of its own

Repairs the shapes of the previous commit's tests:

- commandSegments starts a new command at a word that begins on a new
  line (the grammar puts the newline of a `\cmd` line at the start of its
  word) and drops that newline from the word's value.
- A here-document operator's arguments are only those on its own l…
seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
…ere-documents are data, ps -fu Eve passes

Failed first (tests/test_secret_path_guard.py, run against the guard of 2f097bf): 64 subtest
failures on the new rows, none of them passing by luck. Blocked before and now allowed or
blocked as they should be: 11 BLOCKED rows returned None ("# macOS" then ps -E, "# check ..."
then systemctl --user show-environment, "# run it" then systemd-run --pipe --wait printenv,
echo ${#PATH}; printenv, echo $#; printenv, gh api .../issues/1#c; cat "$PAPER_ENV_FILE",
echo a#b; printenv, two commands with a trailing comment, echo "$(date # )" whose printenv line
follows the comment, a comment holding quotes and a backquote, and a here-document heading
followed by printenv), plus 22 behind rtk proxy and 11 behind a keyring exec. 8 ALLOWED rows
returned a reason: echo ok # "$(printenv)", printf '%s' $'it\'s "$("printenv")"', ps -fu Eve
(two rows) and four commit-message patterns with a quoted here-document. 11 scanner-table rows
and 1 expected-pass-through row (a shell reading a quoted here-document in a substitution)
differed, and all 10 subtests of the five real commit messages (cf58426, 32a6cd2, a84fa7c,
88af2ba, b553826, verbatim, in the standard git commit -m "$(cat <<'EOF' ...)" and the gh pr
create --body forms) failed. The bug behind the first group is older than this work: tokenize()
joins lines with ";" and shlex reads a "#" anywhere, "$#" and "a#b" included, as a comment, so a
"#" dropped the whole rest of the command (found by the 2026-09-29 review).

Design: scan_shell() (substitution_bodies() now calls it) returns the bodies and the comment
spans from the same single pass. A comment is a "#" that starts a word (at the start of the text
or after a blank or one of ;&|()<>, and not right after a quote, escape or substitution), outside
quotes and outside a double-quoted substitution, to the end of the line; inside a $( ) body it
also runs to the end of its line, so a ")" in it closes nothing, and in a backquote body it ends
at the closing backquote. tokenize() removes the spans and sets shlex's commenters to "", and
command_segments() also reads the command as shlex did before (for commands up to 200,000
characters) and adds only the segments that reading lacks, so nothing the guard read before is
dropped. In a here-document body a comment hides the rest of that body, as before, not the text
after the terminator, so a script written through a here-document (its first line is #!) stays as
unread as it was. $'...' is data up to the first quote that no backslash escapes, except inside
double quotes. A here-document whose delimiter is quoted (also partly: <<'E'OF) is data: the body
that scan_shell returns loses it, keeping the operator line and the terminator; an unquoted
delimiter keeps its body, as at the top level. ps_shows_environment() skips the word after a
cluster that ends in a value option (ps -fu Eve). The three vacuous BLOCKED rows are marked as
regression rows and two are replaced by rows that pass at base, echo "$(echo \"$(printenv)\")"
and echo "$(base64 ~/.ssh/id_ed25519)".

Checked: oracle 65 cases, 0 mismatches; suites: only test_host_profile_copy_is_verbatim fails
(139 tests). Differential against base c26800f over 54,126 generated commands: 0 loosened, 0
unexplained reason changes, 0 newly blocked without a new form. Random grammar fuzz against
base, 160,000 strings in 4 seeds: 0 loosened, 0 exceptions. Real repository, by realistic shape:
785 fenced blocks, 4,338 fenced lines and 201 scripts written through a quoted here-document: 0
newly blocked; the synthetic shape of a whole script passed as one command string newly blocks 4
of 201, none of them a true positive and all long-standing readings that the old "#" had
hidden (three array literals x=(env -i ...) and a python set(...)). Commit messages: the
standard pattern newly refuses 0 of 842 (both guards refuse the same 8; the previous tip newly
refused 3 of the 842 and 6 of 1,865 across all refs); the top-level
git commit -F - shape newly blocks the two messages of this work that quote the new ps -E and
systemd-run rules in prose, by those rules and not by the comment change. 1 MB timing: a 840 KB
git commit -m "$(cat <<'EOF' ...)" takes 8.5 s at base and 8.9 s now (two readings would take
17.8 s, so commands above 200,000 characters are read once).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 30, 2026
…ere-documents are data, ps -fu Eve passes

Failed first (tests/test_secret_path_guard.py, run against the guard of 2f097bf): 64 subtest
failures on the new rows, none of them passing by luck. Blocked before and now allowed or
blocked as they should be: 11 BLOCKED rows returned None ("# macOS" then ps -E, "# check ..."
then systemctl --user show-environment, "# run it" then systemd-run --pipe --wait printenv,
echo ${#PATH}; printenv, echo $#; printenv, gh api .../issues/1#c; cat "$PAPER_ENV_FILE",
echo a#b; printenv, two commands with a trailing comment, echo "$(date # )" whose printenv line
follows the comment, a comment holding quotes and a backquote, and a here-document heading
followed by printenv), plus 22 behind rtk proxy and 11 behind a keyring exec. 8 ALLOWED rows
returned a reason: echo ok # "$(printenv)", printf '%s' $'it\'s "$("printenv")"', ps -fu Eve
(two rows) and four commit-message patterns with a quoted here-document. 11 scanner-table rows
and 1 expected-pass-through row (a shell reading a quoted here-document in a substitution)
differed, and all 10 subtests of the five real commit messages (cf58426, 32a6cd2, a84fa7c,
88af2ba, b553826, verbatim, in the standard git commit -m "$(cat <<'EOF' ...)" and the gh pr
create --body forms) failed. The bug behind the first group is older than this work: tokenize()
joins lines with ";" and shlex reads a "#" anywhere, "$#" and "a#b" included, as a comment, so a
"#" dropped the whole rest of the command (found by the 2026-09-29 review).

Design: scan_shell() (substitution_bodies() now calls it) returns the bodies and the comment
spans from the same single pass. A comment is a "#" that starts a word (at the start of the text
or after a blank or one of ;&|()<>, and not right after a quote, escape or substitution), outside
quotes and outside a double-quoted substitution, to the end of the line; inside a $( ) body it
also runs to the end of its line, so a ")" in it closes nothing, and in a backquote body it ends
at the closing backquote. tokenize() removes the spans and sets shlex's commenters to "", and
command_segments() also reads the command as shlex did before (for commands up to 200,000
characters) and adds only the segments that reading lacks, so nothing the guard read before is
dropped. In a here-document body a comment hides the rest of that body, as before, not the text
after the terminator, so a script written through a here-document (its first line is #!) stays as
unread as it was. $'...' is data up to the first quote that no backslash escapes, except inside
double quotes. A here-document whose delimiter is quoted (also partly: <<'E'OF) is data: the body
that scan_shell returns loses it, keeping the operator line and the terminator; an unquoted
delimiter keeps its body, as at the top level. ps_shows_environment() skips the word after a
cluster that ends in a value option (ps -fu Eve). The three vacuous BLOCKED rows are marked as
regression rows and two are replaced by rows that pass at base, echo "$(echo \"$(printenv)\")"
and echo "$(base64 ~/.ssh/id_ed25519)".

Checked: oracle 65 cases, 0 mismatches; suites: only test_host_profile_copy_is_verbatim fails
(139 tests). Differential against base c26800f over 54,126 generated commands: 0 loosened, 0
unexplained reason changes, 0 newly blocked without a new form. Random grammar fuzz against
base, 160,000 strings in 4 seeds: 0 loosened, 0 exceptions. Real repository, by realistic shape:
785 fenced blocks, 4,338 fenced lines and 201 scripts written through a quoted here-document: 0
newly blocked; the synthetic shape of a whole script passed as one command string newly blocks 4
of 201, none of them a true positive and all long-standing readings that the old "#" had
hidden (three array literals x=(env -i ...) and a python set(...)). Commit messages: the
standard pattern newly refuses 0 of 842 (both guards refuse the same 8; the previous tip newly
refused 3 of the 842 and 6 of 1,865 across all refs); the top-level
git commit -F - shape newly blocks the two messages of this work that quote the new ps -E and
systemd-run rules in prose, by those rules and not by the comment change. 1 MB timing: a 840 KB
git commit -m "$(cat <<'EOF' ...)" takes 8.5 s at base and 8.9 s now (two readings would take
17.8 s, so commands above 200,000 characters are read once).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Oct 1, 2026
…ere-documents are data, ps -fu Eve passes

Failed first (tests/test_secret_path_guard.py, run against the guard of 2f097bf): 64 subtest
failures on the new rows, none of them passing by luck. Blocked before and now allowed or
blocked as they should be: 11 BLOCKED rows returned None ("# macOS" then ps -E, "# check ..."
then systemctl --user show-environment, "# run it" then systemd-run --pipe --wait printenv,
echo ${#PATH}; printenv, echo $#; printenv, gh api .../issues/1#c; cat "$PAPER_ENV_FILE",
echo a#b; printenv, two commands with a trailing comment, echo "$(date # )" whose printenv line
follows the comment, a comment holding quotes and a backquote, and a here-document heading
followed by printenv), plus 22 behind rtk proxy and 11 behind a keyring exec. 8 ALLOWED rows
returned a reason: echo ok # "$(printenv)", printf '%s' $'it\'s "$("printenv")"', ps -fu Eve
(two rows) and four commit-message patterns with a quoted here-document. 11 scanner-table rows
and 1 expected-pass-through row (a shell reading a quoted here-document in a substitution)
differed, and all 10 subtests of the five real commit messages (cf58426, 32a6cd2, a84fa7c,
88af2ba, b553826, verbatim, in the standard git commit -m "$(cat <<'EOF' ...)" and the gh pr
create --body forms) failed. The bug behind the first group is older than this work: tokenize()
joins lines with ";" and shlex reads a "#" anywhere, "$#" and "a#b" included, as a comment, so a
"#" dropped the whole rest of the command (found by the 2026-09-29 review).

Design: scan_shell() (substitution_bodies() now calls it) returns the bodies and the comment
spans from the same single pass. A comment is a "#" that starts a word (at the start of the text
or after a blank or one of ;&|()<>, and not right after a quote, escape or substitution), outside
quotes and outside a double-quoted substitution, to the end of the line; inside a $( ) body it
also runs to the end of its line, so a ")" in it closes nothing, and in a backquote body it ends
at the closing backquote. tokenize() removes the spans and sets shlex's commenters to "", and
command_segments() also reads the command as shlex did before (for commands up to 200,000
characters) and adds only the segments that reading lacks, so nothing the guard read before is
dropped. In a here-document body a comment hides the rest of that body, as before, not the text
after the terminator, so a script written through a here-document (its first line is #!) stays as
unread as it was. $'...' is data up to the first quote that no backslash escapes, except inside
double quotes. A here-document whose delimiter is quoted (also partly: <<'E'OF) is data: the body
that scan_shell returns loses it, keeping the operator line and the terminator; an unquoted
delimiter keeps its body, as at the top level. ps_shows_environment() skips the word after a
cluster that ends in a value option (ps -fu Eve). The three vacuous BLOCKED rows are marked as
regression rows and two are replaced by rows that pass at base, echo "$(echo \"$(printenv)\")"
and echo "$(base64 ~/.ssh/id_ed25519)".

Checked: oracle 65 cases, 0 mismatches; suites: only test_host_profile_copy_is_verbatim fails
(139 tests). Differential against base c26800f over 54,126 generated commands: 0 loosened, 0
unexplained reason changes, 0 newly blocked without a new form. Random grammar fuzz against
base, 160,000 strings in 4 seeds: 0 loosened, 0 exceptions. Real repository, by realistic shape:
785 fenced blocks, 4,338 fenced lines and 201 scripts written through a quoted here-document: 0
newly blocked; the synthetic shape of a whole script passed as one command string newly blocks 4
of 201, none of them a true positive and all long-standing readings that the old "#" had
hidden (three array literals x=(env -i ...) and a python set(...)). Commit messages: the
standard pattern newly refuses 0 of 842 (both guards refuse the same 8; the previous tip newly
refused 3 of the 842 and 6 of 1,865 across all refs); the top-level
git commit -F - shape newly blocks the two messages of this work that quote the new ps -E and
systemd-run rules in prose, by those rules and not by the comment change. 1 MB timing: a 840 KB
git commit -m "$(cat <<'EOF' ...)" takes 8.5 s at base and 8.9 s now (two readings would take
17.8 s, so commands above 200,000 characters are read once).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Oct 1, 2026
…ere-documents are data, ps -fu Eve passes

Failed first (tests/test_secret_path_guard.py, run against the guard of 2f097bf): 64 subtest
failures on the new rows, none of them passing by luck. Blocked before and now allowed or
blocked as they should be: 11 BLOCKED rows returned None ("# macOS" then ps -E, "# check ..."
then systemctl --user show-environment, "# run it" then systemd-run --pipe --wait printenv,
echo ${#PATH}; printenv, echo $#; printenv, gh api .../issues/1#c; cat "$PAPER_ENV_FILE",
echo a#b; printenv, two commands with a trailing comment, echo "$(date # )" whose printenv line
follows the comment, a comment holding quotes and a backquote, and a here-document heading
followed by printenv), plus 22 behind rtk proxy and 11 behind a keyring exec. 8 ALLOWED rows
returned a reason: echo ok # "$(printenv)", printf '%s' $'it\'s "$("printenv")"', ps -fu Eve
(two rows) and four commit-message patterns with a quoted here-document. 11 scanner-table rows
and 1 expected-pass-through row (a shell reading a quoted here-document in a substitution)
differed, and all 10 subtests of the five real commit messages (cf58426, 32a6cd2, a84fa7c,
88af2ba, b553826, verbatim, in the standard git commit -m "$(cat <<'EOF' ...)" and the gh pr
create --body forms) failed. The bug behind the first group is older than this work: tokenize()
joins lines with ";" and shlex reads a "#" anywhere, "$#" and "a#b" included, as a comment, so a
"#" dropped the whole rest of the command (found by the 2026-09-29 review).

Design: scan_shell() (substitution_bodies() now calls it) returns the bodies and the comment
spans from the same single pass. A comment is a "#" that starts a word (at the start of the text
or after a blank or one of ;&|()<>, and not right after a quote, escape or substitution), outside
quotes and outside a double-quoted substitution, to the end of the line; inside a $( ) body it
also runs to the end of its line, so a ")" in it closes nothing, and in a backquote body it ends
at the closing backquote. tokenize() removes the spans and sets shlex's commenters to "", and
command_segments() also reads the command as shlex did before (for commands up to 200,000
characters) and adds only the segments that reading lacks, so nothing the guard read before is
dropped. In a here-document body a comment hides the rest of that body, as before, not the text
after the terminator, so a script written through a here-document (its first line is #!) stays as
unread as it was. $'...' is data up to the first quote that no backslash escapes, except inside
double quotes. A here-document whose delimiter is quoted (also partly: <<'E'OF) is data: the body
that scan_shell returns loses it, keeping the operator line and the terminator; an unquoted
delimiter keeps its body, as at the top level. ps_shows_environment() skips the word after a
cluster that ends in a value option (ps -fu Eve). The three vacuous BLOCKED rows are marked as
regression rows and two are replaced by rows that pass at base, echo "$(echo \"$(printenv)\")"
and echo "$(base64 ~/.ssh/id_ed25519)".

Checked: oracle 65 cases, 0 mismatches; suites: only test_host_profile_copy_is_verbatim fails
(139 tests). Differential against base c26800f over 54,126 generated commands: 0 loosened, 0
unexplained reason changes, 0 newly blocked without a new form. Random grammar fuzz against
base, 160,000 strings in 4 seeds: 0 loosened, 0 exceptions. Real repository, by realistic shape:
785 fenced blocks, 4,338 fenced lines and 201 scripts written through a quoted here-document: 0
newly blocked; the synthetic shape of a whole script passed as one command string newly blocks 4
of 201, none of them a true positive and all long-standing readings that the old "#" had
hidden (three array literals x=(env -i ...) and a python set(...)). Commit messages: the
standard pattern newly refuses 0 of 842 (both guards refuse the same 8; the previous tip newly
refused 3 of the 842 and 6 of 1,865 across all refs); the top-level
git commit -F - shape newly blocks the two messages of this work that quote the new ps -E and
systemd-run rules in prose, by those rules and not by the comment change. 1 MB timing: a 840 KB
git commit -m "$(cat <<'EOF' ...)" takes 8.5 s at base and 8.9 s now (two readings would take
17.8 s, so commands above 200,000 characters are read once).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Oct 1, 2026
…ble quotes, macOS ps -E, systemctl show-environment, fail-closed and work budget (#511)

* Guard: read systemd-run as a launcher; block a secret variable on its command line

Failed first: 30 of 31 new BLOCKED rows in tests/test_secret_path_guard.py returned None
instead of their reason (systemd-run --user --pipe --wait cat "$PAPER_ENV_FILE" and its
bash -ic form, printenv and env behind it, an option's value taken for the command, the
-E/--setenv/-p Environment= secret names), with 60 more failures behind rtk proxy and 23
behind a keyring exec; the oracle showed 8 systemd-run MUST_BLOCK mismatches (22 in all).
The 7 new ALLOWED controls (the trading lane's loader path and ordinary units) passed
before and after.

Design: expand() unwraps systemd-run like env and rtk. It skips the launcher's own options
(wrapper_options, from the getopt table of systemd v255 src/run/run.c, plus the value
options of v256-v258 so a newer host still finds its command), keeps the systemd-run
segment in the result, and reads the started command with every rule, a nested bash -ic
string too. segment_reason() blocks a secret variable NAME (SECRET_NAMES) set through
-E/--setenv or -p/--property Environment=, with or without a value, as the new reason
secret_variable_on_command_line: the command line lands in the journal (_CMDLINE) and the
unit's properties travel over the user bus. skip_wrapper_options now delegates to
wrapper_options. The two EXPECTED_PASS_THROUGH rows that recorded the gap moved to BLOCKED.

Checked: oracle 22 -> 14 mismatches (none left for systemd-run); tests.test_secret_path_guard
green except the host-copy comparison, which is fixed by reinstalling the guard after merge.
Differential against b40b3596 over 48,310 generated commands (every table row wrapped in
launchers and suffixes): 0 loosened, 0 reason changes without the new launcher, 0 newly
blocked commands without it. Randomized comparison of the refactored option walker with the
base one: 400,000 cases, 0 mismatches. SHA256SUMS carries the new guard hash; the guard
section of docs/secret-storage.md records the rule.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: read command substitutions inside double quotes as commands

Failed first: 24 of the 26 new BLOCKED rows in tests/test_secret_path_guard.py returned None
instead of their reason (echo "$(printenv)", echo "`printenv`", x="$(printenv)"; echo "$x",
git commit -m "$(printenv)", a backticked `set` in prose inside double quotes, nesting in
either order of quoting, a substitution inside ${x:-...} and $(( ... )), and the credential
file, .env, secret-name, keyring, native-token and trace rules reached through a body), with
48 more failures behind rtk proxy and 23 behind a keyring exec; the oracle showed 6
MUST_BLOCK mismatches for these forms. The 2 rows that already passed are echo "a $(echo
"$(printenv)") b" (the base tokenizer's quote parity happened to expose the inner body) and
a curl row whose reader already reads ~/.ssh. A second red step: 5 of 9 rows of the scanner
table for here-documents failed before the scanner learned to skip their bodies (below).

Design: substitution_bodies() scans the command once with a small frame stack (unquoted,
double-quoted, `$(` body, backquote body) as the Bash Reference Manual describes it: double
quotes keep $ and the backquote special, a backslash escapes only $ ` " and \, single quotes
are data outside double quotes, a `$(` body ends at its matching parenthesis and a backquote
body at the first unescaped backquote. It returns the outermost bodies that start inside
double quotes; expand() reads each one as a command, level by level to 32 levels, so every
rule and every wrapper (rtk, keyring exec, systemd-run, nested shells) applies to it. An
unquoted $( ... ) stays with the tokenizer, and single-quoted text and a backslash-escaped $(
or backquote stay data (the oracle's MUST_ALLOW rows and 12 new ALLOWED rows pin them). The
scanner treats a here-document's body as literal text, as bash's parser does when it looks for
the `)` that ends a `$(`: prose in a commit message (an unbalanced parenthesis, an apostrophe,
backquotes) opens nothing. Without that, one commit message produced 29 spurious bodies. How
the guard reads here-document bodies as commands is untouched (out of scope).

Checked: oracle 14 -> 8 mismatches (the rest belong to ps -E and systemctl). Table test of 30
texts against their expected bodies, and a bounded-nesting test (5000 levels neither raise nor
hang). tests.test_secret_path_guard green except the host-copy comparison. Differential
against b40b3596 over 50,178 generated commands: 0 loosened, 0 reason changes without a new
form. Plain-form equivalence, the promise of this change: for 516 single-line table and oracle
rows in three shapes ("$(ROW)", "`ROW`", X="$(ROW)"), the new guard's verdict on the
double-quoted form is never more lenient than the base guard's verdict on the unquoted
substitution (0 of 1,548) and never differs in reason; it is stricter for the systemd-run rows
(that launcher was unread in both forms) and for one row where shlex fuses `$(>` into one
token in the plain form. Real corpus: 16,089 distinct lines and blocks of this repository's
Markdown fences and shell scripts, base blocks 33, this change 34, 0 loosened; the one new
block is a true positive (env | cut -d= -f1 inside a double-quoted substitution). Real commit
messages, 840 of them in the git commit -m "$(cat <<'EOF' ... EOF)" pattern: the new verdict
equals the base verdict on the unquoted form for 839 (6 -> 8 blocked; the 2 new blocks are a
prose line starting with `set` that the base already flags in the unquoted form, and item 1's
own message quoting cat "$PAPER_ENV_FILE" behind systemd-run).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: block ps -E and dashless ps clusters with a capital E as environment dumps

Failed first: 14 rows in tests/test_secret_path_guard.py returned None instead of
environment_dump (ps -E, -Ewwp 123, -p 123 -E, -A -E, -AE, -eE, -ef -E, -o pid,command -E,
Eww 123, auxE, E, and three keyring-exec forms), with 22 more failures behind rtk proxy and
11 behind a keyring exec; the oracle showed the 6 ps MUST_BLOCK mismatches. The 7 new ALLOWED
controls (ps aux, -ef --sort, -eo pid,etime,args, -o pid,ETIME, -u Eve, -C E, -o etime=)
passed before and after.

Why: macOS documents -E as the environment display, "-E Display the environment as well", and
lists the BSD-style e as "Same as -E" (Apple adv_cmds ps.1, read 2026-09-29). The guard read
only dashless clusters with a lower-case e. A dashed -e is every process on Linux and macOS
("Identical to -A"), so ps -ef stays allowed.

Design: PS_BSD_CLUSTER accepts E as well as e, in a strict superset of the old language, and
ps_shows_environment() reads a dashed word's letters up to the first option that takes a value
(PS_ARG_OPTIONS), so -Ewwp 123 and -p 123 -E are found while -pE, -uE and -u Eve, where the E is
a value, are not. The dashless branch and the value-skipping are unchanged.

Checked: oracle 8 -> 2 mismatches (the two systemctl rows); tests.test_secret_path_guard green
except the host-copy comparison. 20,000 random realistic ps lines against b40b3596: 0 loosened,
5,181 newly blocked (3,899 distinct) and every one carries an E flag; capital-E values and
names still pass.
Differential over 51,174 generated commands: 0 loosened, 0 unexplained reason changes, 0 newly
blocked commands without a new form. Real corpus of 16,089 repository lines: no new block.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: block systemctl show-environment as a service manager's environment dump

Failed first: all 18 new BLOCKED rows in tests/test_secret_path_guard.py returned None instead
of service_manager_environment (systemctl [--user] show-environment, options before and after
the verb, -M/-H and --machine/--host values, `--`, a redirection between options and verb, sudo,
a path to systemctl, timeout, a pipe, bash -c, systemd-run --pipe, and inside a double-quoted
substitution), with 36 more failures behind rtk proxy and 18 behind a keyring exec; the oracle
showed the 2 systemctl MUST_BLOCK mismatches. The 9 new ALLOWED controls (show -p Environment
UNIT, cat, status, list-units, is-active, show -p MainPID --value UNIT, -H host status, restart,
and a unit named show-environment.service) passed before and after.

Why: systemctl(1) 255 says show-environment dumps "the systemd manager environment block. This
is the environment block that is passed to all processes the manager spawns", so any credential
that was imported into that manager lands in the output, as with env.

Design: systemctl_verb() reads systemctl's first word after its options, using the option table
of src/systemctl/systemctl.c at systemd v255 (getopt string "ht:p:P:alqfs:H:M:n:o:iTr.::", no
leading `+`, so options may also follow the verb; the value options are -t -p -P -s -H -M -n -o
and the 25 long ones of v255, plus the v256-v258 additions) through the same option walker as
systemd-run,
with redirections dropped first. segment_reason() returns the new reason
service_manager_environment for the verb show-environment. It is not `environment_dump`, so
the keyring-exec test requires the exact reason, and the rule still applies behind a keyring
exec and rtk proxy. HINTS names the allowed alternative, systemctl show -p Environment UNIT.

Checked: oracle 2 -> 0 mismatches, exit 0 (65 cases); tests.test_secret_path_guard,
tests.test_credential_status, tests.test_credential_tools and tests.test_install_claude_profile
green except the host-copy comparison, which is fixed by reinstalling the guard after merge.
Differential against b40b3596 over 52,672 generated commands: 0 loosened, 0 unexplained reason
changes, 0 newly blocked commands without a new form. Real corpus of 16,089 repository lines: no
new block from this change. Known gap, recorded in docs/secret-storage.md: systemctl show with no
unit prints the manager's own properties, Environment= among them.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: linear-time scan and launcher walk, arithmetic expansion is no command, an internal error blocks

Failed first (tests/test_secret_path_guard.py, run against the previous tip 03d38712): 7 of 9
PATHOLOGICAL timing rows failed, each run in a child process with a 45 s limit and a 3 s bound
on check() itself: 12,000 here-documents 10.6 s, 12,000 distinct delimiters 10.3 s, 12,000
lines of $((1 << 2)) 10.1 s, 20,000 nested env 28.3 s, 20,000 nested rtk proxy and 60,000
nested systemd-run past 45 s, and the 20,000-deep quote nesting returned environment_dump where
the row expected the base verdict (its expected verdict is environment_dump: shlex's quote parity
exposes the innermost command one level down). 10 more rows failed: 8 substitution_bodies rows
for $((...)) and the allowed row env=2; echo "$((env))", and 3 error-handling rows (main()
raised instead of blocking). A hook that times out does not block the call (Claude Code hooks
documentation, "Timeouts", read 2026-09-29), so a slow guard fails open.

Design: substitution_bodies() is one pass over the text with a stack of frames (unquoted,
double-quoted, $( ), backquotes, ${ }, $(( )) ) and a regular expression that jumps to the next
character that matters, instead of stepping through every character. A here-document's terminator is
found with a binary search in a line index built once per text (a rescan per << was the
quadratic part), and an unterminated one is no here-document, as before. $(( is arithmetic
only when a "))" that touches closes it, so $((env)) reads a variable, a << in it is a shift,
real substitutions inside it are still read, and $((printenv) ) is a substitution holding a
subshell. launcher_chain() walks env, rtk and systemd-run hops with an index instead of copying
the rest of the words at each hop (prefix_end() replaces the copying strip_prefix loop); only a
systemd-run hop keeps a segment of its own, because no rule reads an env or rtk hop that starts a
command. expand() stops reading nested bodies at 32 levels or at four times the command's length
plus 64 KiB. main() catches any exception from check() and blocks with one line,
"blocked (guard_error)", without command text or traceback (new HINTS entry); unreadable hook
input still exits 1 and other tools exit 0.

Checked: the nine PATHOLOGICAL rows now take 0.11 to 0.75 s (test bound 3 s); before/after on
this host, seconds: 12,000 here-documents 10.4 -> 0.16 (base 0.14), 12,000 shifts 10.1 -> 0.24
(base 0.22), 6,000 "cat <<X" 2.6 -> 0.11, 20,000 nested env 28.1 -> 0.11 (base 29.1), 20,000
nested rtk 55.1 -> 0.16 (base 55.9), 60,000 nested systemd-run > 60 -> 0.75 (base 0.55, which
did not read it), 1,600 nested systemd-run 9.98 MiB -> 0.23 MiB (base 0.17 MiB), 20,000-deep quote
nesting 4.9 -> 0.72. A realistic 1 MB command is bounded by shlex, unchanged and not part of this
work: a 840 KB git commit -m "$(cat <<'EOF' ...)" takes 8.5 s at base, 8.6 s at 03d38712, 9.0 s
now. Oracle: 65 cases, 0 mismatches. Suites: only test_host_profile_copy_is_verbatim fails.
Differential against base c26800f3 over 52,960 generated commands: 0 loosened, 0 unexplained
reason changes, 0 newly blocked without a new form; against 03d38712 the only differences are
65 $((...)) forms that pass again (base passed them). Walker comparison with base, 400,000 random
cases: 0 mismatches. Random grammar fuzz, 120,000 strings against base: 0 loosened, 0 exceptions.
Real repository corpus (16,119 lines and blocks): base blocks 33, now 34, the same single true
positive as before, 0 loosened. tests: check() runs without raising over 834 rows: BLOCKED 303, KEYRING_BLOCKED 127,
ALLOWED 119, SAFE_CORPUS 139, EXPECTED_PASS_THROUGH 36, the four oracle groups (now copied into the
test file) 65, the 36 scanner-table texts and the 9 pathological inputs.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: a comment hides only its own line, ANSI-C strings and quoted here-documents are data, ps -fu Eve passes

Failed first (tests/test_secret_path_guard.py, run against the guard of 2f097bf4): 64 subtest
failures on the new rows, none of them passing by luck. Blocked before and now allowed or
blocked as they should be: 11 BLOCKED rows returned None ("# macOS" then ps -E, "# check ..."
then systemctl --user show-environment, "# run it" then systemd-run --pipe --wait printenv,
echo ${#PATH}; printenv, echo $#; printenv, gh api .../issues/1#c; cat "$PAPER_ENV_FILE",
echo a#b; printenv, two commands with a trailing comment, echo "$(date # )" whose printenv line
follows the comment, a comment holding quotes and a backquote, and a here-document heading
followed by printenv), plus 22 behind rtk proxy and 11 behind a keyring exec. 8 ALLOWED rows
returned a reason: echo ok # "$(printenv)", printf '%s' $'it\'s "$("printenv")"', ps -fu Eve
(two rows) and four commit-message patterns with a quoted here-document. 11 scanner-table rows
and 1 expected-pass-through row (a shell reading a quoted here-document in a substitution)
differed, and all 10 subtests of the five real commit messages (cf584265, 32a6cd2b, a84fa7c4,
88af2baa, b5538261, verbatim, in the standard git commit -m "$(cat <<'EOF' ...)" and the gh pr
create --body forms) failed. The bug behind the first group is older than this work: tokenize()
joins lines with ";" and shlex reads a "#" anywhere, "$#" and "a#b" included, as a comment, so a
"#" dropped the whole rest of the command (found by the 2026-09-29 review).

Design: scan_shell() (substitution_bodies() now calls it) returns the bodies and the comment
spans from the same single pass. A comment is a "#" that starts a word (at the start of the text
or after a blank or one of ;&|()<>, and not right after a quote, escape or substitution), outside
quotes and outside a double-quoted substitution, to the end of the line; inside a $( ) body it
also runs to the end of its line, so a ")" in it closes nothing, and in a backquote body it ends
at the closing backquote. tokenize() removes the spans and sets shlex's commenters to "", and
command_segments() also reads the command as shlex did before (for commands up to 200,000
characters) and adds only the segments that reading lacks, so nothing the guard read before is
dropped. In a here-document body a comment hides the rest of that body, as before, not the text
after the terminator, so a script written through a here-document (its first line is #!) stays as
unread as it was. $'...' is data up to the first quote that no backslash escapes, except inside
double quotes. A here-document whose delimiter is quoted (also partly: <<'E'OF) is data: the body
that scan_shell returns loses it, keeping the operator line and the terminator; an unquoted
delimiter keeps its body, as at the top level. ps_shows_environment() skips the word after a
cluster that ends in a value option (ps -fu Eve). The three vacuous BLOCKED rows are marked as
regression rows and two are replaced by rows that pass at base, echo "$(echo \"$(printenv)\")"
and echo "$(base64 ~/.ssh/id_ed25519)".

Checked: oracle 65 cases, 0 mismatches; suites: only test_host_profile_copy_is_verbatim fails
(139 tests). Differential against base c26800f3 over 54,126 generated commands: 0 loosened, 0
unexplained reason changes, 0 newly blocked without a new form. Random grammar fuzz against
base, 160,000 strings in 4 seeds: 0 loosened, 0 exceptions. Real repository, by realistic shape:
785 fenced blocks, 4,338 fenced lines and 201 scripts written through a quoted here-document: 0
newly blocked; the synthetic shape of a whole script passed as one command string newly blocks 4
of 201, none of them a true positive and all long-standing readings that the old "#" had
hidden (three array literals x=(env -i ...) and a python set(...)). Commit messages: the
standard pattern newly refuses 0 of 842 (both guards refuse the same 8; the previous tip newly
refused 3 of the 842 and 6 of 1,865 across all refs); the top-level
git commit -F - shape newly blocks the two messages of this work that quote the new ps -E and
systemd-run rules in prose, by those rules and not by the comment change. 1 MB timing: a 840 KB
git commit -m "$(cat <<'EOF' ...)" takes 8.5 s at base and 8.9 s now (two readings would take
17.8 s, so commands above 200,000 characters are read once).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: read what the review found unread (launchers, $(<FILE), here-documents, systemctl show, quoted wrappers), correct the claims

Failed first (tests/test_secret_path_guard.py, run against the guard of 0ad4835c): 154 subtest
failures on the new rows, each a reason expected and None returned. 41 BLOCKED and KEYRING rows,
plus 78 behind rtk proxy and 31 behind a keyring exec (numbers before the leading-redirection
rows, which failed the same way): a backquote in a single-quoted string
handed to bash -c, sh -c and eval; redirections between systemd launcher options
(systemd-run --user 2>/tmp/log --pipe printenv, ... > /tmp/log -E HF_TOKEN true); $(<.env),
$(< ~/.aws/credentials), x=$(<~/.netrc), echo `<.env`; unquoted here-document bodies that run
$(...) (cat <<EOF with value: "$(printenv)"); run0, systemd-inhibit and systemd-cat; a bare
systemctl show and show -p Environment; /usr/bin/sudo, /usr/bin/timeout, /usr/bin/nice; ps -CE
and ps -C -E; systemctl and systemd-run behind watch and flock inside a keyring exec (the wrapper
matrix rows). 3 scanner-table rows for here-document bodies failed too. The 19 new ALLOWED
controls (the alternatives the docs and the hint recommend, $(<version.txt), a quoted
here-document, path-qualified launchers of ordinary commands, ps -C python3, systemctl show
naming a unit) and the 12 new expected-pass-through rows passed before and after, as they must.

Design: scan_shell() also returns the backquotes that sit in single-quoted and ANSI-C strings
(tokenize() keeps them, so the shell that string is handed to sees them: item 9), the bodies of
an unquoted here-document scanned as double-quoted text with no closing quote (a "hd" frame, one
substring scan per body, nested at most 8 deep), and the quoted here-document bodies inside a
substitution that closes, which the tokenizer no longer reads in either of its readings (the
five real commit messages, and two commit messages of this work that the earlier guard refused
through shlex's quote parity, now pass in the standard pattern). wrapper_options() skips
redirections for the systemd launchers only: for timeout the duration in "timeout 5 > out cmd"
reads as a descriptor and the command was swallowed (the base differential caught it), so every
other wrapper keeps the walk it had. A token "(<" (shlex joins punctuation that touches) and a
segment that is only "< FILE" add a "cat FILE" segment, and a leading "< FILE cmd" adds "cmd < FILE",
in addition to what was read before.
The systemd launcher tables (run0 from parse_argv_sudo_mode v256-v262, systemd-inhibit,
systemd-cat) and the value options of later systemd-run releases cite the tag that first has
each (--root-directory v259, --output v261: the review found them absent from v256-v258).
systemctl_call() reads the verb and its arguments wherever the options stand, so a bare show
(or show -p Environment) with no unit is refused. prefix_end() compares the program name of a
wrapper, so path-qualified launchers count. ps_shows_environment() refuses an E word right
after -C. launched_commands() reads systemctl and the systemd launchers, and keyring_reason()
returns service_manager_environment for a started systemctl. The earlier reading of a command
that holds a # (or a protected backquote, or data) is kept beside the new one, for commands up
to 200,000 characters, so nothing the guard read before is dropped. The dated docs subsection is
rewritten with the corrected claims: the manager's "Started <unit> - <command line>" journal
line (measured on systemd 255.4, not _CMDLINE); --pipe returns the output, --wait shows terse
unit information and the output goes to the journal without either; the scope of "every rule
applies"; the alternatives considered for reading shell syntax (bashlex, tree-sitter-bash,
mvdan/sh, with the dates and licences read from PyPI and the GitHub API on 2026-09-29); the
residuals with their inert strings; and the systemctl hint now names the alternatives. Three
files that recommended the refused command (omniroute.service, token-report-refresh.service,
tools/token-report/README.md) keep it and add that it is for the operator's own terminal, with
the agent alternative.

Checked: oracle 65 cases, 0 mismatches; suites (4 named): only test_host_profile_copy_is_verbatim
fails (141 tests); the suites that read the three edited files (omniroute unit, token-report
units, ecosystem manifest, dashboard data, workflow hardening, adoption docs and status, 319
tests) pass. Differential against base c26800f3 over 59,530 generated commands: 0 loosened, 0
unexplained reason changes, 0 newly blocked without a new form. Random grammar fuzz against base,
240,000 strings in 6 seeds: 0 loosened, 0 exceptions. Mutation fuzz of strings the base blocks
(quotes, comments, backquotes, here-documents and launchers inserted), 8 seeds of 22,440
mutants: 1 differs (the run before the last change, seeds 21 to 28, had 3), every one a quoted
here-document whose body holds the printenv (bash runs nothing there). Walker comparison with base, 400,000 random cases: 0 mismatches. Real repository
by shape: 785 fenced blocks, 4,338 fenced lines and 201 scripts written through a quoted
here-document newly block 0; a whole script passed as one command string newly blocks 4 of 201
(array literals x=(env -i ...) and a python set(...), read as commands as before). Commit
messages: the standard pattern newly refuses 0 of 1,884 (all refs); two messages of this work that
the earlier guard refused there now pass; the unquoted-delimiter and the top-level git commit -F -
shapes read prose as commands as before (newly refused 8 and 4 of 843, by the # fix and by this
work's own new rules quoting themselves in prose). tests: check() runs without raising over 968
rows.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: a quoted here-document is data only for a substitution whose command takes data

The data rule of 172596ed treated the body of a quoted here-document inside a double-quoted
command substitution as data whatever received the substitution. The coordinator's probes found
the cost: bash re-reads what a substitution prints as code for `eval "$(...)"`, `bash -c
"$(...)"` (also sh, zsh, behind env, timeout, nohup, xargs) and a code consumer nested in an
echo's substitution, and the base guard refused all of those as environment_dump while 172596ed
allowed them.

Failed first (the new tests against 172596ed): 196 failures and 2 errors, one failure being the
tolerated host-copy comparison. The 38 new BLOCKED rows failed directly (38), behind rtk proxy
(76) and behind a keyring exec (38); 27 consumer rows, 13 substitution-body rows and 2 oracle
groups failed; the two data-rule tests errored on the missing names.

Design: scan_shell() decides each quoted here-document only after the whole text is read.
data_consumers() reads the text once more with every returned substitution replaced by a marker
and the comments and top-level here-document bodies cut, finds the command each marker is an
argument of after the reserved words, assignments, wrappers and launchers the guard models
(receiving_command), and answers True only for git, gh, echo, printf, cat and tee. A shell, eval,
an interpreter, source, xargs, watch, ssh or any other program, an assignment, the command
position, a marker before the program word, a substitution nested in another, and any text the
guard cannot tell (markers in the text, more than 100,000 characters left after the cuts, more
than 16 launcher hops) keep the old reading: the body read as command lines and the tokenizer
input whole. Only a here-document at the command level of the returned body is data, so
`echo "$(echo "$(cat <<'EOF' ...)")"` keeps the strict reading. cat and tee count only for a
here-string word (`cat <<< "$(...)"`): a matrix of 39 heads x 18 launcher prefixes x 5
here-document forms x 8 payloads x 4 substitution forms (112,320 commands) showed the plain
allowlist loosening `cat "$(cat <<'EOF' ... a credential path ...)"`, which the base refused
through the reader rules, because a reader's operand is a file name; the operand forms now keep
the base verdicts.

Checked: kw/guard_probe_final.py base candidate reports loosened rows: 0 (172596ed: 1); the
eight-consumer probe, with rtk proxy and keyring exec forms, 0 of 13 loosened (172596ed: 13);
oracle 65 cases, 0 mismatches; the matrix loosens 0 rows for every head that is no data
consumer (172596ed: up to 2,040 per head) and 0 for git, gh, echo and printf; mutation fuzz, 60
mutants per each of 397 blocked strings: 6 loosened at seed 7 and 2 at seed 11 (172596ed: 387
and 378), each a mutant whose broken terminator or quote makes the rest of a git or echo
substitution here-document data that bash never runs, or a bash syntax error; differential
against c26800f3 over 59,592 commands from 936 rows: 0 loosened; the real-repository corpora
(fence lines and blocks, scripts through here-documents): 0 loosened; tests.test_secret_path_guard
29 tests, the host-copy comparison the only failure.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: main() refuses a command of more than 600,000 characters (command_too_large)

Claude Code's hooks documentation ("Timeouts", read 2026-09-29) says a timed-out command hook
does not block the tool call, and this hook's timeout is 10 s. Measured on this host on
2026-09-29 with one quoted word: 1,000,000 characters took 10.2 to 13.0 s in check() (base
c26800f3 and this version alike, over several runs), 600,000 took 3.8 to 4.4 s and 500,000 took
2.9 to 3.3 s, so a large enough command outlasts the hook and runs unread.

Failed first: test_the_size_limit_is_at_600000_characters and
test_a_command_over_the_size_limit_is_refused_without_echoing_it errored on the missing
MAX_COMMAND_CHARACTERS and the missing reason. The third test, a 500,000-character ordinary
command that must not be refused for its size, passed before and after, as it should.

Design: main() refuses len(command) > MAX_COMMAND_CHARACTERS (600,000 characters, not bytes)
before any rule reads the command, with exit 2 and one stderr line that names no command text:
"blocked (command_too_large)" and the hint to put the content in a file with the Write tool and
pass the path. check() is unchanged for size, so every table row behaves as before, and a
command of exactly 600,000 characters still reaches check().

Checked: a 700,000-character command is refused in under 5 s through the real hook subprocess
(exit 2, one line, neither the marker text nor its filler on stderr or stdout); 500,000
characters of ordinary shell (quoted words, arithmetic, redirections, separators) exit 0 with no
output inside 5 s and, with printenv appended, exit 2 with environment_dump; the boundary
test passes 600,000 characters to check() (mocked) and refuses 600,001 without calling it; 300,000
four-byte characters, 1.2 MB of UTF-8, pass. tests.test_secret_path_guard: 32 tests, the
host-copy comparison the only failure.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard docs: the data-consumer rule and the size limit, SHA256SUMS carries the new guard

docs/secret-storage.md, guard section, dated 2026-09-29 and no restructuring: one new item for
the rule that a quoted here-document is data only for git, gh, echo, printf, cat (as a
here-string word) and tee. It says why (bash runs what a substitution prints for a shell with
-c, eval, source, xargs or an interpreter), that every unlisted consumer keeps the stricter
reading so prose in a quoted here-document inside them can still be refused, what the guard
still does not read (an allowed consumer's output used later, a shell reading the quoted
here-document itself), and the measured matrix. The here-document item now says "where the
command takes data", and the timeout item gains one sentence for command_too_large with the
measured timings (1,000,000 characters of one quoted word: 10.2 to 13.0 s; 600,000: 3.8 to
4.4 s; 500,000: 2.9 to 3.3 s), replacing the "not fixed" note that described the gap. The
guard's own docstring says the same in two sentences.

adoption/hooks/claude/SHA256SUMS: only the guard's line changed, to
914342a040d9632b9aedc435e58617a78808b5f5f6fcbe88a9460edda41be75c (91,845 bytes); sha256sum -c
over the file reports every line OK.

Checked: kw/guard_oracle.py exits 0 (65 cases, 0 mismatches); kw/guard_probe_final.py reports
loosened rows: 0 with the timing rows of the last round unchanged (12,000 here-documents 0.19 s,
20,000-deep quoting 0.66 s, 60,000 systemd-run 0.58 s); tests.test_secret_path_guard,
test_credential_status, test_credential_tools and test_install_claude_profile ran 148 tests, the
host-copy comparison the only failure; 305 further tests (docs consistency, adoption status,
blind checkout, effort guard, release pins, landscape) pass; scripts/validate.py lists only the
manifest hash and byte mismatches of the pinned files this branch changed (SHA256SUMS, the
guard, its tests, docs/secret-storage.md, and the two files 172596ed changed without a re-pin,
omniroute.service and tools/token-report/README.md).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: CANARY_E2E_KEY is a secret name (the variable of the canary-e2e inventory entry)

A later inventory entry, canary-e2e (a disposable synthetic proof key: class test_canary, status
test_only), declares the variable CANARY_E2E_KEY, and test_inventory_secret_names_are_all_guarded
requires every inventory name to be in SECRET_NAMES, so it would fail there without the name.
It is added now, ahead of that entry (which belongs to another branch and is not touched here),
and is refused like every other name.

Failed first (the two new rows against 914342a0, the guard of 82767cae): 9 failures, one being the
tolerated host-copy comparison; the rows failed directly (2), behind rtk proxy (4) and behind a
keyring exec (2). The four forms `echo "$CANARY_E2E_KEY"`, `echo $CANARY_E2E_KEY`, the
os.environ lookup and `rg -n CANARY_E2E_KEY` all returned None; they now return
secret_variable_reference (three) and secret_name_search, and `rg -n MY_CANARY_E2E_KEY_HINT`
still passes (the word boundary).

Design: one entry in SECRET_NAMES with a comment; the expansion, lookup and search rules and
test_every_secret_name_is_caught_by_a_search pick it up from the tuple. Explicit BLOCKED rows for
`echo "$CANARY_E2E_KEY"` and `rg -n CANARY_E2E_KEY`, so the wrapper tests also run them behind rtk
proxy and a keyring exec. No docs sentence lists where the names come from (only the comment above
the tuple does), so the docs are unchanged. The inventory tie holds with and without the later
entry's variable (29 inventory names, 30 unique guard names, none missing). SHA256SUMS carries the
final guard hash for this and the earlier changes of this round.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: refuse a command of more than 200,000 characters, not 600,000

The independent verification review asked for a limit low enough that the tokenizer's two readings
of a command (with and without shlex's comments) always both run, so that no reading is skipped
above some length. Measured on this host on 2026-09-29, one quoted word costs 0.6 s at 200,000
characters (about 1.5 s for both readings), 2.9 to 3.3 s at 500,000, 3.8 to 4.4 s at 600,000 and
10.2 to 13.0 s at 1,000,000 (base c26800f3 and this version alike), against the hook's 10 s.

Failed first: the size test that pins the limit failed (600000 != 200000) once it asked for
200,000; the others use the constant and follow it. Tests: a 250,000-character command is
refused through the real hook (exit 2, one line, no command text) inside 5 s; a 150,000-character
ordinary command exits 0 with no output inside 5 s and, with printenv appended, exits 2 with
environment_dump; the boundary test passes 200,000 characters to check() (mocked) and refuses
200,001 without calling it; 150,000 four-byte characters (600 KB of UTF-8) pass.

The constant's comment, the guard docstring and the docs sentence carry the new limit and the
timings; SHA256SUMS carries the guard's new hash (a stale one fails four tests of the other
suites).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: the dashless ps cluster test is a set-membership test, not a backtracking regular expression

The independent verification review found that `'ps ' + 'E' * 70000 + 'q; printenv'` took 13.2 s in
check() against 0.08 s at the base guard: PS_BSD_CLUSTER, `^[aAcfhjlmrsStTuvwxXLnE]*[eE][aAcefhjlmrsStTuvwxXLnE]*$`,
has two overlapping quantifiers that fail in quadratic time, and a hook that outlasts its 10 s
timeout does not block the call. Reproduced here at 15.2 s.

Failed first: the new timing row "ps cluster of 70,000 E and a letter that is no flag" failed at
15.2 s against the 3 s bound, and the equivalence test errored on the missing function.

Design: is_ps_bsd_cluster(word) is PS_BSD_LETTERS.issuperset(word) and an `e` or `E` in it, the
same language as the regular expression (checked on every word of up to 5 letters over nine
letters, 66,430 words, and on 13 words the guard meets). Five timing rows for ps clusters of
70,000 characters (a lead of E, of e, of a, a trailing E, a letter that is no flag) join the
PATHOLOGICAL table with their verdicts and the 3 s bound, and a test times every compiled pattern
of the guard (module level, the scan and store tables: 52) on 70,000 repeats of 12 characters
with a lead and a tail that make a match fail late, each under a quarter of a second. Negative
control: the same test against the guard before this change fails naming PS_BSD_CLUSTER at 14.6 s.
The audit script that found nothing else (PAREN_INPUT is slow under search, 6.4 s, but the guard
only calls it as fullmatch, which is linear) also timed 2,132 check() shapes, a first word
(ps, env, sudo, systemd-run, systemctl, keyctl, rtk, grep, git, bash -c ...) followed by 70,000
repeats of a character or a token: the slowest were the two ps shapes, everything else 0.3 s.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: no general here-document reading; a strict canonical idiom is the only exemption

The independent verification review of 172596ed (GPT-6, effort max, read-only) found six blocking
block-to-allow regressions, five in the general here-document machinery: a quoted delimiter made
the body of a here-document inside a double-quoted substitution data, and each imperfection of
the bash-exact parsing that needs erased executable text. Reproduced here as inert strings, each
an environment_dump at the base guard c26800f3 and None at 172596ed: `echo "$(bash <<'EOF'` with
`echo "$(printenv)"` in its body (a shell reads the body); a `<<$'EOF'` delimiter; a body line
`text\` that the guard's backslash-newline join glued to its terminator; an arithmetic command
`((1 << "2"))` taken for a `<<`; a quoted `#` (`printf ' #x'`) taken for a comment that hid the
rest of a here-document body. The design change of the coordinator replaces the general reading.

Failed first (the new tests against 37132506): 125 failures and 129 errors, one failure being the
tolerated host-copy comparison: the recognizer table (61 rows), the consumer table (60 rows), the
41 new BLOCKED rows directly (29), behind rtk proxy (58) and behind a keyring exec (29), the
plain-scan rows of the bodies table (6), the ANSI-C rows (7), a known-bypass row, the neutral
reading and the joined-line test.

Design. scan_shell() no longer reads here-documents at all: the delimiter regular expressions,
heredoc_delimiter, heredoc_end, the terminator index in the scanner, the here-document data
spans, the comment-in-a-body rule and the `hd` frame are deleted, so a here-document's lines are
command lines as the tokenizer has always read them at the top level (a `#` hides its own line,
a `"$(...)"` in a body line is a substitution). The one exemption is idiom_spans(): a
double-quoted word that is exactly `"$(` blanks `cat` blanks `<<` [`-`] blanks `'IDENT'` blanks
newline, the body, the FIRST line that is exactly IDENT (for `<<-` after its leading tabs, no
other trimming, no carriage return), blanks and newlines, `)"`, standing alone as a word. It is
looked for in the text as written, before check() joins backslash-newline pairs, with a line
index built once and a tail memoised per terminator, so no head costs a rescan. exempt_idioms()
then reads the text once more with each idiom replaced by a marker and asks the command that
receives it: git, gh, echo or printf after the reserved words, assignments, wrappers and
launchers the guard models, with the idiom a word of its own after the program, not inside the
body of a substitution, and the text readable by shlex without its fallback. Then neutral_reading()
puts a neutral word in its place for the COMMAND reading only; the raw-text rules (store paths,
secret names and expansions, /proc) read the whole text as before. Everything else, a shell,
eval, an interpreter, source, xargs, watch, ssh, cat, tee, an assignment, the command position,
`\EOF`, `"EOF"`, `$'EOF'` and unquoted delimiters, text on the operator line, text after the
terminator, an unterminated body and a word glued to `--body=`, keeps the old reading, the body
lines read as commands. The legacy-reading cutoff is deleted with the machinery. shlex knows no
ANSI-C quoting, so tokenize() keeps each `$'...'` string as one word (harmless
`printf '%s' $'it\'s #\nprintenv\n'` passes; `bash -c $'printenv'` is read as `bash -c printenv`).

Checked at this state: the oracle 65 cases, 0 mismatches; kw/guard_probe_final.py base candidate
`loosened rows: 0`; the eight consumer variants and their rtk proxy and keyring exec forms 0 of 13
loosened; the reviewer's five here-document strings and the coordinator's eight consumer forms
are BLOCKED rows and environment_dump; the four suites 151 tests, the host-copy comparison the
only failure; 14 mutation controls of the recognizer, the consumer decision and the tokenizer
change each make a test fail (two survivors found by the first run, a plain `<<` terminator with
a leading tab and the second tokenizer reading, got a row each). The launcher-redirection strings
and the corrections of the review's other findings follow in their own commits.

Residual gaps, recorded as EXPECTED_PASS_THROUGH rows: a git or gh option that runs its value
(`git rebase --exec "$(cat <<'EOF' ...)"`, `gh alias set`), an echo piped into a shell or written
into a script that runs later, and an unquoted here-document that expands `$(...)` between single
quotes (the base guard passed it too). Friction, in the strict direction: prose in a body that
starts a line with `printenv` is refused outside the idiom, and so are `--body="$(...)"`, `cat`
and `tee` as consumers and a text with an unbalanced apostrophe before the idiom.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: a redirection between launcher hops is read wherever it stands

The review of 172596ed found that `env -u < "$PAPER_ENV_FILE" UNUSED cat` and `systemd-run --pipe
--unit < "$PAPER_ENV_FILE" demo cat` are credential_file_read at the base guard and None at the
tip: the index-based launcher walk took the redirection operator for the value of `-u` and of
`--unit` and dropped the segment that held the credential-file operand. Reproduced with a
matrix of redirection positions: for env, env -i, systemd-run, sudo, timeout, timeout -s, nice,
rtk proxy, a keyring exec and four chains (sudo env, env sudo, nohup timeout), every position
between the hops, each of `<`, `<<<`, `2>`, `>`, `>>` and `&>`, to a credential pointer, `.env`
and `~/.aws/credentials`, 684 combinations: 8 rows looser than the base guard, all `<` or `<<<`
to the pointer in an option-value slot (`env -u < ...`, `systemd-run ... --unit < ...`, the same
behind sudo and env).

Failed first (the new tests against the guard of the previous commit): 170 failures, one being
the tolerated host-copy comparison: 152 subtests of the position matrix, 13 reviewer strings, and
the wrapper tests behind rtk proxy (2) and a keyring exec (1) and the direct table (1).

Design, following the coordinator's instruction: the raw segment is read as written beside every
launcher walk (command_segments appends it before any walk), so a redirection is checked
wherever a walk would put it; and no option-value slot swallows a redirection any more.
skip_redirections() steps over a redirection operator and its target, and wrapper_options,
env_command_start, rtk_command_start and prefix_end use it wherever they look for an option, a
value or the started command (only an operator token starts one: a bare number may be a value,
`nice -n 5 > out cmd`). An input redirection that was stepped over is appended to the command that
gets it, as if it stood after it (launcher_chain now returns it), so `env -u UNUSED < .env cat` is
read as `cat < .env`, and a keyring exec and a leading here-string do the same. Nothing is
carried for a number before an operator: `sudo 2>/dev/null -u root printenv` stays a recorded gap.

Checked at this state: for every launcher, position and operator the verdict equals the verdict of
the same command with the redirection written last, 684 combinations, 0 mismatches; against the
base guard 0 rows are looser and 166 are stricter, each the operand of a `cat` that the base
guard only saw when the redirection stood after the command; kw/guard_probe_final.py and the
oracle unchanged (`loosened rows: 0`, 65 cases, 0 mismatches); the smoke row of each reviewer
string is credential_file_read; the four suites 153 tests, the host-copy comparison the only
failure; 22 mutation controls (the earlier 14 and eight for the walkers, the raw segment and the
carried redirections) each make a test fail. `nice > /tmp/out -n 5 cat .env`, a recorded gap of
172596ed, is now a BLOCKED row. The redirection rows also cover a launcher with no command
(`sudo < "$PAPER_ENV_FILE"`), which only the raw segment sees.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: correct where a systemd-run value goes, and document the two loosenings and the ps friction

The review's other findings, each a correction of a claim, not of a rule.

systemd-run provenance. The guard's hint, docstring, comments and docs said that the manager logs
systemd-run's command line, an `-E NAME=value` included, in the journal. Read from upstream
(systemd v255, src/run/run.c, fetched 2026-09-29): the default Description, which the manager
logs as `Started <unit> - <description>`, is `quote_command_line(arg_cmdline)` (lines 1940 to
1951), and arg_cmdline is the command and its arguments after the options; `-E` fills
arg_environment (line 348), which is appended to the start message as the unit's Environment
property (lines 853 to 866). So a value given with `-E`, `--setenv` or `-p Environment=` is in
the command line of the systemd-run process, which the process listing shows while it runs, and
in the transient unit's Environment property on the user bus (`systemctl --user show -p
Environment UNIT`, and any bus client); it is not in the journal, while a value written after the
command is (a recorded gap). The rule, a secret name on a systemd-run line is refused, is
unchanged. The verbatim commit message of cf584265 that the tests carry as a real message keeps
its old wording, as it is a historical text.

Failed first: test_the_systemd_run_hint_says_where_a_value_on_the_command_line_goes failed against
the previous guard (the hint named the journal); the rest of this commit is documentation and
rows that the previous guard already satisfied (ALLOWED rows for `ps -fu steve`, `ps -fu eve` and
`ps -u steve` pass there too).

The list of loosenings. The base guard c26800f3 refuses `ps -fu steve` and `ps -fu eve`, because
it reads the user name as a dashless BSD cluster with an `e` (`ps -u steve` it passes); `-u`
takes a value, and this work lets them pass. Together with the canonical idiom of the
here-document item, that is the only loosening against the base guard: docs/secret-storage.md now
says so in one item and the ALLOWED rows carry it. `ps -CEmacs` stays refused, documented as
friction (on procps `-C` takes a command name: `ps -C emacs`).

Checked: the oracle 65 cases, 0 mismatches; kw/guard_probe_final.py `loosened rows: 0`; the four
suites 154 tests, the host-copy comparison the only failure.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: an idiom inside any command substitution is never exempt; corrected measurements

My own adversarial probe of the exemption before the final pass, nested and wrapping constructs
against the base guard, found that `eval $(echo "$(cat <<'EOF'` + `printenv` + `EOF` + `)")` passed
(the base guard passed it too, so it is no regression, but bash runs `printenv` there): the
tokenizer splits an unquoted `$(echo IDENT)` into a segment of its own whose command is echo, and
exempt_idioms() judged the idiom by that segment, while the substitution's output goes to eval.
The same held for `bash -c $(echo IDIOM)`, `bash <<< $(echo IDIOM)`, `x=$(echo IDIOM); eval $x`,
`xargs sh -c $(...)`, `python3 -c $(...)`, `eval $(git log -1 --format=%B IDIOM)`, a subshell or
brace group around any of them, and `echo $(git commit -m IDIOM)`. The dq-substitution rule of the
previous commit covered only the substitutions inside double quotes.

Failed first (the new tests against the guard of the previous commit): 68 failures and 16 errors,
one failure the tolerated host-copy comparison: 15 rows of the exemption table, the 14 new
BLOCKED rows directly (13), behind rtk proxy (26) and behind a keyring exec (13), and the two
tests of the new scan result (10 and 6 errors).

Design: scan_shell() now returns a fifth result, the (start, end) of the text inside each
outermost command substitution, `$(...)` or a backquote pair, quoted or not (a counter of open
substitution frames, kept right through the `$((printenv) )` conversion and an unterminated
substitution); exempt_idioms() refuses any idiom whose marker lies inside one, which replaces the
bodies-only rule. A subshell `( ... )` and a brace group are no substitution: `( cd sub && git
commit -m IDIOM )` stays exempt. The probe's 28 rows: 0 looser than the base guard; every nested
form above is now an environment_dump; `git commit -m IDIOM` and `echo IDIOM` pass; the pipe into a
shell stays a recorded gap.

Docs: the exemption item names the rule (`eval $(echo "$(cat <<'EOF' ... EOF)")`), and the measured
numbers of the item are corrected to the final ones: the consumer matrix is 40 consumers x 18
launcher prefixes x 5 here-document forms x 8 payloads x 4 substitution forms = 115,200 commands
(the item said 39 and 112,320), the commit history is 1,956 messages (13 refused by the base guard,
10 by the guard, all by raw-text rules, 4 that the base refused pass), the differential over 60,493
derived commands loosens 102 rows and every one is a `ps -fu steve` or `ps -fu eve`, and the
mutation fuzz of 24,840 mutants of 414 blocked strings loosens 38 to 44 per seed, of which 36 to 41
are variants of those two rows and 2 or 3 are mutants whose broken terminator leaves a real quoted
body (the item said 24,720 mutants, 412 strings and 2 or 3, without the ps rows).

Checked: the oracle 65 cases, 0 mismatches; kw/guard_probe_final.py `loosened rows: 0`; the four
suites 155 tests, the host-copy comparison the only failure; 22 mutation controls each make a test
fail (the rule that this commit strengthens is killed by the new rows).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: the docstring says what the exemption is; the reviewer's launcher strings are table rows

Housekeeping the design change left: the guard's module docstring still described the superseded
rule ("a quoted here-document inside a substitution whose command is git, gh, echo, printf, cat or
tee"), and said backslash-newline is joined first without saying that the canonical idiom is looked
for before that. It now says: a here-document's lines are command lines like any other; the one
exemption is a strict canonical idiom, a double-quoted word of exactly the shape
`"$(cat <<'EOF'` newline, body, `EOF` newline, `)"`, behind git, gh, echo or printf, whose body is
data for the command reading while the raw-text rules still read it. Three test comments that
named "data consumers" or "the first repair round" follow.

The two launcher strings of the review, `env -u < "$PAPER_ENV_FILE" UNUSED cat` and `systemd-run
--pipe --unit < "$PAPER_ENV_FILE" demo cat`, were checked by a test that asserts only that they
are refused; the coordinator asked for them in the tables, so they are BLOCKED rows
(credential_file_read, as at the base guard), where the wrapper tests also run them behind rtk
proxy and a keyring exec.

No behaviour changes: the four suites and the oracle pass as before, and the guard's hash in
SHA256SUMS moves with the docstring.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: a redirection is no argument of set, export, declare and typeset

A grammar fuzz of my own, launcher chains with redirections at random places and idioms in random
placements, generated against the base guard (10 seeds of 60,000 commands: 5 with the launcher
walkers of the previous commits), found 4 commands in 120,000 that the base guard blocks and the
guard allowed: a keyring exec with a redirection between the variable name and the `--`, for
example `python3 scripts/kernel_keyring.py exec tavily_api_key TAVILY_API_KEY < x.txt -- set`
(base: environment_dump_in_keyring_exec). The base guard dropped that redirection when it took the
command after `--`; the walkers of this work carry a skipped input redirection to the command it
belongs to, and is_environment_dump() counted the carried `< x.txt` as arguments of `set`
(`set a b` sets positional parameters), so the dump was no dump. `set < FILE` and `set > FILE`
still print every variable.

Failed first (the new rows against the guard of the previous commit): 53 failures, one the
tolerated host-copy comparison: 16 BLOCKED rows directly, 28 behind rtk proxy, 8 behind a keyring
exec. Four of the new rows contain a keyring exec of their own and live in KEYRING_BLOCKED, which the
wrapper tests do not prefix with another one.

Design: the module docstring already says "redirection operands are never taken for arguments";
is_environment_dump() now reads set, export, declare and typeset through command_arguments(), which
drops the redirections, so `set > /tmp/vars`, `set < /dev/null`, `export -p > FILE`, `declare -p >
FILE`, `typeset -p 2> err` and the keyring forms are an environment_dump, while `set -e >
/dev/null`, `export FOO=1 > /dev/null` and `declare -a items > /dev/null` (a real argument) pass.
This only tightens: the base guard passed the redirected forms.

Checked: 10 seeds of the grammar fuzz, 600,000 commands, 0 looser than the base guard (seeds 1 and 2,
which found the four, included); the four suites and the oracle unchanged; the docs sentence is in
the launcher-redirection item.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard docs: the measurements of the here-document item at the final tip

The here-document item of the guard section still quoted the numbers of an earlier commit of this
series (1,956 commit messages, 13 and 10 refused, 4 passing, a differential over 60,493 commands, mutation
fuzz of 24,840 mutants of 414 strings). Replaced with the numbers measured at 660e6812, the tip whose
guard is byte-identical to this one: 1,960 messages, 10 refused (base 14), 5 that the base guard refused
now pass, a differential over 62,365 derived commands that loosens 102 (all `ps -fu steve` or
`ps -fu eve`), a grammar fuzz of 10 seeds of 60,000 launcher chains that loosens none, and mutation fuzz
of 25,140 mutants of 419 blocked strings that loosens 33 to 44 per seed (29 to 38 `ps` variants, 4 to 6
mutants whose broken terminator leaves a real quoted body that bash prints and does not run). The
loosenings item says 5 messages at 660e6812 instead of 4. Two over-long lines of the item re-wrapped.

Docs only: the guard is unchanged (sha256 72d43a2a15c4ccc002a13d9747ed818ee51832084811753912d4e9aa4ebf347e,
94,067 bytes), so the SHA256SUMS line stays. The four suites: 155 tests, one failure, the tolerated
host-copy comparison (the host profile copy is not reinstalled), 3 skipped.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard docs: the commit-message count is over all refs, not one commit's history

The measurement of the here-document item said "1,960 commit messages reachable there" at 660e6812.
The script that counts them (git log --all) reads every ref of the shared repository, including other
branches and worktrees, so the number is the distinct messages on all refs when measured (1,964 on
the run that found this) and grows without any change to the guard. The item now says that, names the
guard as the one of 660e6812, and the loosenings item points at that measurement instead of repeating
"at 660e6812". Re-measured on the current refs: 10 messages refused by the guard, 14 by the base guard,
5 that the base guard refused now pass (3899d49b, 2f097bf4, 172596ed, c3bdbaf5, 5b964612), so the
sentences hold.

The timeout item claimed that both readings of a command always run inside the hook's 10 s at the
200,000-character limit. It now carries the measurement behind that: seven inputs of about 199,000
characters that make every reading run (double-quoted text with a `#` throughout plus one idiom or
1,000 idioms, the same single-quoted, ANSI-C strings, one quoted word, many backquotes, many subshells)
took 0.2 to 3.3 s in check() on this host, the slowest being the double-quoted text plus one idiom.
Three over-long lines (the loosenings item and two in the timeout item) re-wrapped.

Docs only: the guard is unchanged (sha256 72d43a2a15c4ccc002a13d9747ed818ee51832084811753912d4e9aa4ebf347e,
94,067 bytes) and the SHA256SUMS line stays.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: remove the canonical-idiom exemption; a here-document body is read as command lines behind every consumer

The second verification review (GPT-6 through the gateway lane, of 50ca6ca2) found six more blocking
regressions, five of them inside the strict canonical-idiom exemption that replaced the general
here-document reading: an idiom head inside another here-document's body that swallowed its
terminator, an idiom inside an unquoted here-document that a shell reads, an echo piped into a shell,
`git rebase --exec` (a git option that runs its value), a case pattern's `)` that closed a
substitution early, and a process substitution that runs a shell. Each is an environment_dump at the
base guard and None at that tip. The coordinator's decision: the exemption is a net risk for its
friction benefit, so it goes, and the guard is tightening-only again. Two reviews and twelve
block-to-allow regressions show that telling a data consumer from a code consumer by the words of one
command line is what fails.

Failed first (the six reviewer strings as BLOCKED rows against 50ca6ca2): 24 failures, 6 direct, 12
behind rtk proxy and 6 behind a keyring exec; the two `ps -Ccat e` strings of the same review are the
next commit.

Design: idiom_spans, exempt_idioms, receiving_command, neutral_reading and everything only they used
(IDIOM_HEAD, IDIOM_TAIL, IDIOM_FOLLOWERS, IDIOM_CONSUMERS, RESERVED_STARTERS, SUBSTITUTION_MARKER,
MAX_LAUNCH_HOPS, the bisect import, and scan_shell's fifth result with its open-substitution counter)
are deleted. check() reads the joined command as it did before the idiom existed: a here-document's
lines are command lines, quoted delimiter or not, behind git, gh, echo and printf as behind a shell,
eval, an interpreter, source, xargs, watch and ssh. The raw-text rules, the double-quoted
substitution reading, comments, the launcher redirection handling, the systemd-run family, ps -E,
systemctl and CANARY_E2E_KEY stay.

Tests: the six strings and 16 documented-friction rows are BLOCKED rows (a commit message or pull-request
body whose prose has a line that reads as a dump; a literal `$(printenv)` in a quoted here-document
in file-writing, Python, commit-message and PR-comment workflows; Python's `set()` after a comment
line in an interpreter's here-document), ALLOWED has the workarounds (`git commit -F`, `gh ... --body-file`)
and a prose message that passes; the five idiom tests and their tables are gone, the real commit
messages of this repository are recorded as the friction they now are (each refused with its reason),
and a matrix of 26 consumers, 7 delimiter forms and 10 dumps (1,820 inert strings) must all be refused,
so an exemption that returns for one consumer fails there.

Checked: oracle phase base 65/65; the probe with the base guard `loosened rows: 0`; the four suites 151
tests, the one failure the tolerated host-copy comparison; every reviewer string of both reviews is
blocked. Of 2,015 distinct commit messages on all refs, in the `git commit -m` and `gh pr create --body`
patterns, the guard refuses 36 and the base guard 16, and none that the base guard refuses passes. The
seven worst-case shapes of about 199,000 characters now take at most 1.0 s (3.3 s with the third
reading). The guard section of docs/secret-storage.md states the friction and the workaround, and
its loosenings item says one; SHA256SUMS carries the new guard hash.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: ps is read as both hosts read it (procps and macOS); `-C` has two readings and either one refuses

The second verification review of 50ca6ca2 found that removing `-C` from the value-taking options of a
cluster lost the procps reading: `ps -Ccat e` and `ps -fCcat e` (environment_dump at the base guard)
returned None. On procps `-C cmdlist` takes a command name (`cat`) and the BSD `e` after it shows the
environment; on macOS `-C` is a flag, so `-Ccat` is `-C -c -a -t` and `t` takes `e` as a tty. The
coordinator's instruction: keep both readings and refuse when either shows the environment.

Failed first (the new rows against 752def7f): 16 failures, 4 direct (`ps -Ccat e`, `-fCcat e`,
`-Ccat eww`, `-Ccat E`), 8 behind rtk proxy and 4 behind a keyring exec. `ps -Cnginx e` and `-fCnginx
auxe` were refused before by chance (the `g` of `nginx` is a value letter) and stay refused.

Design: ps_reading_shows_environment(words, procps) is one host's reading, ps_shows_environment() the
union. procps (PS_ARG_OPTIONS, `-C` a value option, dashless BSD words anywhere, no dashed `-E`) and
macOS (PS_MACOS_ARG_OPTIONS, `-C` a flag, dashed `-E`, a dashless option string only as the first
argument) each keep their own state of which word a value option takes. The `-C -E` special case of the
hybrid reading is gone: the macOS reading finds it. Sources: procps-ng 4.0.4 ps(1) on this host (`-C
cmdlist`, `e` "Show the environment after the command", no `-E`), and Apple adv_cmds ps/ps.c at
60bc9ebf (fetched 2026-09-29): PS_ARGS `aACcdeEfg:G:hjLlMmO:o:p:rSTt:U:u:vwx`, `case 'C': rawcpu = 1`,
`kludge_oldps_options(..., argv[1], ...)` applied to argv[1] only, and a later non-option word is a
process id or an "illegal argument". Checked on this host: `ps -Cbash u` honours the BSD `u` after
the glued command name and `ps -fC bash u` reports conflicting format options, so a dashless word
after a UNIX option is a BSD option on procps. `ps -CEmacs` stays refused (friction: the macOS
reading; write `ps -C emacs`).

Loosening, the one this work makes, now stated generally: a value after a clustered value-taking
option is no BSD flag cluster (`ps -fu steve`, `-fo user`, `-ft e`, `-fU steve`, `-fC e`); the base
guard skipped the value after a stand-alone `-u` but not after `-fu`. A differential over 478,915 ps
command lines (every one of up to three words from a vocabulary of 65, plus 200,000 random ones of
four to six words) loosens 6,673 against the base guard, every one holding such a cluster with its
value (0 outside that family); the reference two-host model of the tests (a per-letter state machine)
and the guard agree on all 478,915. The tests carry that reference over 160,434 lines (a vocabulary
of 54), ALLOWED rows for the value forms and BLOCKED rows for the union.

Checked: oracle 65/65; probe with the base guard `loosened rows: 0`; the four suites 152 tests, the one
failure the tolerated host-copy comparison. The docs replace the ps item and the loosenings item, and
record macOS's legacy mode (`-e` read as `-E` when ps.c's `u03` is off) as a gap. SHA256SUMS carries the
new guard hash.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Guard: one work budget per check() call; a command that needs more is refused as command_too_complex

The second verification review found two commands far below the 200,000-character cap that outlast the
hook's 10 s timeout, where a timed-out PreToolUse hook blocks nothing: a keyring exec with 1,200
`python3` words (9,645 characters; the analysis read every suffix of them as a command, 5.8 million
characters through shlex: base 2.1 to 2.8 s, 50ca6ca2 12 to 13 s) and 195,000 emoji in five nested
double-quoted substitutions (195,068 characters; four-byte characters cost four times as much to
tokenize and the text was read at every level: base 1.6 s, 50ca6ca2 15.7 to 17.9 s here, over 25 s
in the review). My own scan found the same class elsewhere: 4,000 distinct keyring execs (25 s), 3,500
keyring execs of one variable, each starting a distinct program (268 s: a quadratic test of mentions
against spans and a scan of the whole text for each started command), a chain of 6,000 nested keyring
execs (a copy of the rest of the words at every hop, past 30 s), and a keyring exec whose started
command is a 180,000-character argument (6.4 s).

Failed first (the new tests against d46f6c0c): the six unit tests 2 failures and 10 errors (no
WorkBudgetExceeded, WORK_LIMITS, storage_width or start_work, expand kept 3 copies of one segment,
no command_too_complex hint); the timing rows in the tests' own child process: 1,200 words 12.48 s,
600 words 3.16 s, 5,000 words past the 45 s limit, the emoji input 17.20 s, 120,000 euro signs 11.0 s.

Design: check() starts a budget, reads the command, stops it (start_work, stop_work; outside a check
call nothing is counted). WORK_LIMITS: (a) characters passed to shlex over every reading and nesting
level, each at its storage width (1 byte ASCII and Latin-1, 2 for the rest of the BMP, 4 beyond: shlex
costs 0.41, 0.85 and 1.70 s on 200,000 characters of one quoted word), 400,000; (b) texts read 10,000,
words emitted (each segment its words plus one) 1,000,000, launched-command reads of the keyring
analysis 500. lex() charges before shlex runs, command_segments() charges each text and each emitted
segme…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant