Skip to content

Command guard tightening: systemd-run launchers, substitutions in double quotes, macOS ps -E, systemctl show-environment, fail-closed and work budget - #511

Merged
seathatflowsinourveins merged 35 commits into
mainfrom
claude/guard-launchers-20260929
Oct 1, 2026
Merged

seathatflowsinourveins merged 35 commits into
mainfrom
claude/guard-launchers-20260929

Conversation

@seathatflowsinourveins

@seathatflowsinourveins seathatflowsinourveins commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Scope

  • What this PR changes, in one or two sentences: tightens the host-wide Claude Code command guard (scripts/hooks/secret_path_guard.py). Changes:

    • systemd-run, run0, systemd-inhibit and systemd-cat are read as launchers. Before this change a reader they start was not inspected: systemd-run --user --pipe --wait cat "$PAPER_ENV_FILE" returned the pointed-to file.
    • $(...) and backticks inside double quotes are read as commands.
    • Comments hide only their own line.
    • macOS ps -E forms and systemctl show-environment (and a bare systemctl show) are refused.
    • An internal error and a command over 200,000 characters exit 2; a timed-out hook does not block the call.
    • One shared work budget per check() call bounds tokenisation and launcher analysis (command_too_complex).
    • Every command is also read as the installed guard reads it, and is refused when either reading refuses, so no command the installed guard refuses is allowed.
    • Four store-path patterns that backtracked (they timed out the hook at 10 s) are linear scans.
    • CANARY_E2E_KEY is a secret name.

    There is no loosening.

  • Base commit: 6bbef3c6 (main at the rebase; head dc33b48a)

  • Lane: lane:foundation

  • Owned paths touched: scripts/hooks/secret_path_guard.py, tests/test_secret_path_guard.py, adoption/hooks/claude/SHA256SUMS (the guard's pin line) and the guard section of docs/secret-storage.md. Three doc notes recommended a command the guard now refuses: adoption/templates/systemd/omniroute.service line 26, adoption/templates/systemd/token-report-refresh.service and tools/token-report/README.md. The change there is additive: the operator's command stays, with the agent-safe alternative added. Also the receipt evidence/receipts/guard-k3-verification-20260930.json and the last commit's manifests/evidence.json. Not touched: the host copy (reinstalled after merge and after the Gate A window closes, announced first), any credential file.

  • Not in this PR, stated plainly: the next guard change (K4) carries the following, from a written contract with its own review:

    • the guard model of the key runner;
    • the interpreter here-document false positive (set(a) | set(b) in a python3 - body);
    • the OmniRoute management routes (/api/ on the gateway ports, with an allowlist; decision record "OmniRoute management API (decision, 2026-09-30)");
    • user-manager environment writes;
    • a keyring store from a literal;
    • CLAUDE_CODE_OAUTH_TOKEN;
    • procps personality selectors;
    • the two residuals below.

SOTA sources

  • systemd v255 (systemd/systemd@db11bab3):
    • src/run/run.c and src/systemctl/systemctl.c option tables: value options and boolean flags; -E NAME[=VALUE], -p Environment=.
    • The journal Description is built from the started command's arguments after the options (run.c:1940-1951).
    • -E NAME=value puts the value in the caller's argv and in the transient unit's Environment property; -E NAME forwards the caller's value into the property (env-util.c:417, run.c:348, run.c:853-866).
    • systemd-run(1), systemctl(1) and journalctl(1) 255 man pages, plus the value options of v256 to v258.
  • ps on both hosts:
    • Apple adv_cmds ps/ps.c at 60bc9ebf (PS_ARGS, case 'C', kludge_oldps_options applied to argv[1] only) and ps(1).
    • procps-ng 4.0.4 man page (-C takes a command name; a dashless BSD word is read anywhere; PS_PERSONALITY and I_WANT_A_BROKEN_PS select BSD parsing).
    • ps is refused when either host's reading shows the environment.
  • Bash Reference Manual: Command Substitution, Double Quotes, Comments, Here Documents and Redirections. A number is a descriptor only when it touches the operator.
  • Claude Code hooks documentation, PreToolUse and Timeouts (code.claude.com/docs/en/hooks, read 2026-09-29): only exit 2 blocks, and a timed-out command hook does not block the call. That is why the guard fails closed on an internal error, refuses oversized or over-budget commands, and replaces backtracking patterns with linear scans.
  • Python re documentation (the module backtracks; there is no linear-time guarantee): the four affected patterns were rewritten as explicit scans, with an old-pattern-versus-scan equivalence test.

Evidence-class table

Claim Evidence class Command / receipt
No command the installed guard refuses is allowed by this guard local_integration All four checks below report no such command. Receipt guard-k3-verification-20260930.
— acceptance oracle local_integration guard_oracle.py: 65 cases, 0 mismatches
— probes local_integration side-by-side probe and launcher option-value probe: 0 loosened rows each
— documented-command replay local_integration 1,490 documented commands: 0 newly blocked, 0 loosened, 0 reason changed
— builder harnesses local_integration 460,910-command equivalence of the second reading with the installed guard; differential over 88,207 commands; grammar and mutation fuzz; 15 of 15 mutation controls killed
Slow inputs no longer time out the hook local_integration Real hook process on the third review's inputs: 0.06-0.14 s, exit 2. The installed guard timed out at 10 s on three and took 4.03 s on one. Receipt guard-k3-verification-20260930.
Independent reviews source_review plus probes by the reviewer Two read-only GPT-6 verifications (gpt-6-astra, effort max, OmniRoute gateway lane), one before and one after the repair round (below)

Local commands run

$ python3 -m unittest tests.test_secret_path_guard
Ran 46 tests ... FAILED (failures=1)   # test_host_profile_copy_is_verbatim only: the installed host copy stays the old guard until reinstall; CI skips it
$ python3 -B guard_oracle.py <checkout>
summary: 65 cases, 0 mismatches (phase base)
$ python3 -B guard_diff_corpus.py <installed guard> <this guard> <checkout>
replayed 1490 distinct documented commands (1758 lines)
newly blocked: 0; loosened: 0; reason changed: 0
$ python3 scripts/validate.py
{"components": 69, "hashed_files": 8297, "profiles": 4, "receipts": 175, "status": "passed"}

Review and repair record

  1. Earlier designs. A general here-document data rule, then a canonical git commit idiom exemption. Each was rejected after an independent GPT-6 review found six blocking block-to-allow regressions in it. Both are removed.
  2. Verification 3 (GPT-6, of the pre-repair head b00ac4e1):
    • All 15 earlier blocking strings were fixed.
    • New findings:
      • blocking: a procps personality selector (PS_PERSONALITY=old ps -axu e) passed the clustered-value ps loosening;
      • high: four store-path patterns were quadratic before any work accounting (inherited from the installed guard; 11.9-15 s);
      • medium: a find -ok prefix walk (3.9 s);
      • low: set 1 > out newly refused;
      • notes.
    • My own probe added four systemd-run --description kernel_keyring.py exec ... block-to-allow rows. The repair found 11 more of the same class.
  3. One repair round.
    • The ps loosening is removed.
    • The second reading (the installed guard's walk, ported function by function) makes "no loosening" hold by construction.
    • Linear scans for the four store-path patterns; find prefixes walked by index; descriptor adjacency; docs corrected.
  4. Verification 4 (GPT-6, of the repaired guard, same bytes as this head):
    • Every verification 3 finding is fixed.
    • No command the installed guard refuses is allowed.
  5. Residuals, carried to K4 (the user allows one repair round):
    • Verification 4 labels one finding blocking against the unmerged pre-repair candidate: segment deduplication drops a descriptor-bearing duplicate, so set 0 < /dev/null; set 0</dev/null passes. The installed guard allows it too, so this PR does not loosen anything.
    • A documentation note: the adjacency statement omits the conservative fallback for a command containing a \x01 byte.
  6. Friction, all in the stricter direction:
    • A commit message or PR body written through "$(cat <<'EOF' ... EOF)" whose prose has a line that reads as an environment dump is refused. Write it with the editor tool and pass git commit -F <file> or gh ... --body-file <file>.
    • ps -fu steve and ps -fo user are refused again, as by the installed guard. ps -f -U 1000 and pgrep -u steve pass.
    • More than 1,410 bare env hops, or 997 two-word hops, is command_too_complex.

Decision record

The guard section of docs/secret-storage.md and docs/decisions/2026-09-29-key-management.md: the status section's item 1 and the dated 2026-09-30 sections.

Host evidence

Not applicable to this PR: no file under evidence/hosts/ changes.

Merge and reinstall happen only after the Gate A owner (native-agent-stack-2d) says window W is closed: the frozen check hooks.carriers_match_repo compares the installed hooks with main. Then, announced first:

  1. gh pr merge 511 --squash --match-head-commit <head>.
  2. cp -p ~/.claude/hooks/secret_path_guard.py ~/.claude/hooks/secret_path_guard.py.bak-<date>.
  3. python3 tools/adoption/install_claude_profile.py --only guard from the live clone.
  4. Read back the sha256 against the pin, then run tests.test_secret_path_guard.
  5. Send the Gate A owner the hash.

Checklist

  • New/changed GitHub Actions are pinned to a full commit SHA with a version comment (no floating tags): no workflow changes.
  • New/changed workflows declare top-level permissions: contents: read: no workflow changes.
  • No secrets are printed, logged or committed; no new required secret was added without a documented owner. All guard tests use inert strings.
  • No new paid hosting, subscription or billing surface was introduced.
  • Peer-owned untracked files and worktrees were preserved (not deleted, moved or overwritten).

🤖 Generated with Claude Code

@seathatflowsinourveins seathatflowsinourveins added the lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers label Sep 29, 2026
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, you can upgrade your account or add credits to your account and enable them for code reviews in your settings.

Scout and others added 23 commits October 1, 2026 01:47
… command line

Failed first: 30 of 31 new BLOCKED rows in tests/test_secret_path_guard.py returned None
instead of their reason (systemd-run --user --pipe --wait cat "$PAPER_ENV_FILE" and its
bash -ic form, printenv and env behind it, an option's value taken for the command, the
-E/--setenv/-p Environment= secret names), with 60 more failures behind rtk proxy and 23
behind a keyring exec; the oracle showed 8 systemd-run MUST_BLOCK mismatches (22 in all).
The 7 new ALLOWED controls (the trading lane's loader path and ordinary units) passed
before and after.

Design: expand() unwraps systemd-run like env and rtk. It skips the launcher's own options
(wrapper_options, from the getopt table of systemd v255 src/run/run.c, plus the value
options of v256-v258 so a newer host still finds its command), keeps the systemd-run
segment in the result, and reads the started command with every rule, a nested bash -ic
string too. segment_reason() blocks a secret variable NAME (SECRET_NAMES) set through
-E/--setenv or -p/--property Environment=, with or without a value, as the new reason
secret_variable_on_command_line: the command line lands in the journal (_CMDLINE) and the
unit's properties travel over the user bus. skip_wrapper_options now delegates to
wrapper_options. The two EXPECTED_PASS_THROUGH rows that recorded the gap moved to BLOCKED.

Checked: oracle 22 -> 14 mismatches (none left for systemd-run); tests.test_secret_path_guard
green except the host-copy comparison, which is fixed by reinstalling the guard after merge.
Differential against b40b359 over 48,310 generated commands (every table row wrapped in
launchers and suffixes): 0 loosened, 0 reason changes without the new launcher, 0 newly
blocked commands without it. Randomized comparison of the refactored option walker with the
base one: 400,000 cases, 0 mismatches. SHA256SUMS carries the new guard hash; the guard
section of docs/secret-storage.md records the rule.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Failed first: 24 of the 26 new BLOCKED rows in tests/test_secret_path_guard.py returned None
instead of their reason (echo "$(printenv)", echo "`printenv`", x="$(printenv)"; echo "$x",
git commit -m "$(printenv)", a backticked `set` in prose inside double quotes, nesting in
either order of quoting, a substitution inside ${x:-...} and $(( ... )), and the credential
file, .env, secret-name, keyring, native-token and trace rules reached through a body), with
48 more failures behind rtk proxy and 23 behind a keyring exec; the oracle showed 6
MUST_BLOCK mismatches for these forms. The 2 rows that already passed are echo "a $(echo
"$(printenv)") b" (the base tokenizer's quote parity happened to expose the inner body) and
a curl row whose reader already reads ~/.ssh. A second red step: 5 of 9 rows of the scanner
table for here-documents failed before the scanner learned to skip their bodies (below).

Design: substitution_bodies() scans the command once with a small frame stack (unquoted,
double-quoted, `$(` body, backquote body) as the Bash Reference Manual describes it: double
quotes keep $ and the backquote special, a backslash escapes only $ ` " and \, single quotes
are data outside double quotes, a `$(` body ends at its matching parenthesis and a backquote
body at the first unescaped backquote. It returns the outermost bodies that start inside
double quotes; expand() reads each one as a command, level by level to 32 levels, so every
rule and every wrapper (rtk, keyring exec, systemd-run, nested shells) applies to it. An
unquoted $( ... ) stays with the tokenizer, and single-quoted text and a backslash-escaped $(
or backquote stay data (the oracle's MUST_ALLOW rows and 12 new ALLOWED rows pin them). The
scanner treats a here-document's body as literal text, as bash's parser does when it looks for
the `)` that ends a `$(`: prose in a commit message (an unbalanced parenthesis, an apostrophe,
backquotes) opens nothing. Without that, one commit message produced 29 spurious bodies. How
the guard reads here-document bodies as commands is untouched (out of scope).

Checked: oracle 14 -> 8 mismatches (the rest belong to ps -E and systemctl). Table test of 30
texts against their expected bodies, and a bounded-nesting test (5000 levels neither raise nor
hang). tests.test_secret_path_guard green except the host-copy comparison. Differential
against b40b359 over 50,178 generated commands: 0 loosened, 0 reason changes without a new
form. Plain-form equivalence, the promise of this change: for 516 single-line table and oracle
rows in three shapes ("$(ROW)", "`ROW`", X="$(ROW)"), the new guard's verdict on the
double-quoted form is never more lenient than the base guard's verdict on the unquoted
substitution (0 of 1,548) and never differs in reason; it is stricter for the systemd-run rows
(that launcher was unread in both forms) and for one row where shlex fuses `$(>` into one
token in the plain form. Real corpus: 16,089 distinct lines and blocks of this repository's
Markdown fences and shell scripts, base blocks 33, this change 34, 0 loosened; the one new
block is a true positive (env | cut -d= -f1 inside a double-quoted substitution). Real commit
messages, 840 of them in the git commit -m "$(cat <<'EOF' ... EOF)" pattern: the new verdict
equals the base verdict on the unquoted form for 839 (6 -> 8 blocked; the 2 new blocks are a
prose line starting with `set` that the base already flags in the unquoted form, and item 1's
own message quoting cat "$PAPER_ENV_FILE" behind systemd-run).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…onment dumps

Failed first: 14 rows in tests/test_secret_path_guard.py returned None instead of
environment_dump (ps -E, -Ewwp 123, -p 123 -E, -A -E, -AE, -eE, -ef -E, -o pid,command -E,
Eww 123, auxE, E, and three keyring-exec forms), with 22 more failures behind rtk proxy and
11 behind a keyring exec; the oracle showed the 6 ps MUST_BLOCK mismatches. The 7 new ALLOWED
controls (ps aux, -ef --sort, -eo pid,etime,args, -o pid,ETIME, -u Eve, -C E, -o etime=)
passed before and after.

Why: macOS documents -E as the environment display, "-E Display the environment as well", and
lists the BSD-style e as "Same as -E" (Apple adv_cmds ps.1, read 2026-09-29). The guard read
only dashless clusters with a lower-case e. A dashed -e is every process on Linux and macOS
("Identical to -A"), so ps -ef stays allowed.

Design: PS_BSD_CLUSTER accepts E as well as e, in a strict superset of the old language, and
ps_shows_environment() reads a dashed word's letters up to the first option that takes a value
(PS_ARG_OPTIONS), so -Ewwp 123 and -p 123 -E are found while -pE, -uE and -u Eve, where the E is
a value, are not. The dashless branch and the value-skipping are unchanged.

Checked: oracle 8 -> 2 mismatches (the two systemctl rows); tests.test_secret_path_guard green
except the host-copy comparison. 20,000 random realistic ps lines against b40b359: 0 loosened,
5,181 newly blocked (3,899 distinct) and every one carries an E flag; capital-E values and
names still pass.
Differential over 51,174 generated commands: 0 loosened, 0 unexplained reason changes, 0 newly
blocked commands without a new form. Real corpus of 16,089 repository lines: no new block.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…nment dump

Failed first: all 18 new BLOCKED rows in tests/test_secret_path_guard.py returned None instead
of service_manager_environment (systemctl [--user] show-environment, options before and after
the verb, -M/-H and --machine/--host values, `--`, a redirection between options and verb, sudo,
a path to systemctl, timeout, a pipe, bash -c, systemd-run --pipe, and inside a double-quoted
substitution), with 36 more failures behind rtk proxy and 18 behind a keyring exec; the oracle
showed the 2 systemctl MUST_BLOCK mismatches. The 9 new ALLOWED controls (show -p Environment
UNIT, cat, status, list-units, is-active, show -p MainPID --value UNIT, -H host status, restart,
and a unit named show-environment.service) passed before and after.

Why: systemctl(1) 255 says show-environment dumps "the systemd manager environment block. This
is the environment block that is passed to all processes the manager spawns", so any credential
that was imported into that manager lands in the output, as with env.

Design: systemctl_verb() reads systemctl's first word after its options, using the option table
of src/systemctl/systemctl.c at systemd v255 (getopt string "ht:p:P:alqfs:H:M:n:o:iTr.::", no
leading `+`, so options may also follow the verb; the value options are -t -p -P -s -H -M -n -o
and the 25 long ones of v255, plus the v256-v258 additions) through the same option walker as
systemd-run,
with redirections dropped first. segment_reason() returns the new reason
service_manager_environment for the verb show-environment. It is not `environment_dump`, so
the keyring-exec test requires the exact reason, and the rule still applies behind a keyring
exec and rtk proxy. HINTS names the allowed alternative, systemctl show -p Environment UNIT.

Checked: oracle 2 -> 0 mismatches, exit 0 (65 cases); tests.test_secret_path_guard,
tests.test_credential_status, tests.test_credential_tools and tests.test_install_claude_profile
green except the host-copy comparison, which is fixed by reinstalling the guard after merge.
Differential against b40b359 over 52,672 generated commands: 0 loosened, 0 unexplained reason
changes, 0 newly blocked commands without a new form. Real corpus of 16,089 repository lines: no
new block from this change. Known gap, recorded in docs/secret-storage.md: systemctl show with no
unit prints the manager's own properties, Environment= among them.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… command, an internal error blocks

Failed first (tests/test_secret_path_guard.py, run against the previous tip 03d3871): 7 of 9
PATHOLOGICAL timing rows failed, each run in a child process with a 45 s limit and a 3 s bound
on check() itself: 12,000 here-documents 10.6 s, 12,000 distinct delimiters 10.3 s, 12,000
lines of $((1 << 2)) 10.1 s, 20,000 nested env 28.3 s, 20,000 nested rtk proxy and 60,000
nested systemd-run past 45 s, and the 20,000-deep quote nesting returned environment_dump where
the row expected the base verdict (its expected verdict is environment_dump: shlex's quote parity
exposes the innermost command one level down). 10 more rows failed: 8 substitution_bodies rows
for $((...)) and the allowed row env=2; echo "$((env))", and 3 error-handling rows (main()
raised instead of blocking). A hook that times out does not block the call (Claude Code hooks
documentation, "Timeouts", read 2026-09-29), so a slow guard fails open.

Design: substitution_bodies() is one pass over the text with a stack of frames (unquoted,
double-quoted, $( ), backquotes, ${ }, $(( )) ) and a regular expression that jumps to the next
character that matters, instead of stepping through every character. A here-document's terminator is
found with a binary search in a line index built once per text (a rescan per << was the
quadratic part), and an unterminated one is no here-document, as before. $(( is arithmetic
only when a "))" that touches closes it, so $((env)) reads a variable, a << in it is a shift,
real substitutions inside it are still read, and $((printenv) ) is a substitution holding a
subshell. launcher_chain() walks env, rtk and systemd-run hops with an index instead of copying
the rest of the words at each hop (prefix_end() replaces the copying strip_prefix loop); only a
systemd-run hop keeps a segment of its own, because no rule reads an env or rtk hop that starts a
command. expand() stops reading nested bodies at 32 levels or at four times the command's length
plus 64 KiB. main() catches any exception from check() and blocks with one line,
"blocked (guard_error)", without command text or traceback (new HINTS entry); unreadable hook
input still exits 1 and other tools exit 0.

Checked: the nine PATHOLOGICAL rows now take 0.11 to 0.75 s (test bound 3 s); before/after on
this host, seconds: 12,000 here-documents 10.4 -> 0.16 (base 0.14), 12,000 shifts 10.1 -> 0.24
(base 0.22), 6,000 "cat <<X" 2.6 -> 0.11, 20,000 nested env 28.1 -> 0.11 (base 29.1), 20,000
nested rtk 55.1 -> 0.16 (base 55.9), 60,000 nested systemd-run > 60 -> 0.75 (base 0.55, which
did not read it), 1,600 nested systemd-run 9.98 MiB -> 0.23 MiB (base 0.17 MiB), 20,000-deep quote
nesting 4.9 -> 0.72. A realistic 1 MB command is bounded by shlex, unchanged and not part of this
work: a 840 KB git commit -m "$(cat <<'EOF' ...)" takes 8.5 s at base, 8.6 s at 03d3871, 9.0 s
now. Oracle: 65 cases, 0 mismatches. Suites: only test_host_profile_copy_is_verbatim fails.
Differential against base c26800f over 52,960 generated commands: 0 loosened, 0 unexplained
reason changes, 0 newly blocked without a new form; against 03d3871 the only differences are
65 $((...)) forms that pass again (base passed them). Walker comparison with base, 400,000 random
cases: 0 mismatches. Random grammar fuzz, 120,000 strings against base: 0 loosened, 0 exceptions.
Real repository corpus (16,119 lines and blocks): base blocks 33, now 34, the same single true
positive as before, 0 loosened. tests: check() runs without raising over 834 rows: BLOCKED 303, KEYRING_BLOCKED 127,
ALLOWED 119, SAFE_CORPUS 139, EXPECTED_PASS_THROUGH 36, the four oracle groups (now copied into the
test file) 65, the 36 scanner-table texts and the 9 pathological inputs.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ere-documents are data, ps -fu Eve passes

Failed first (tests/test_secret_path_guard.py, run against the guard of 2f097bf): 64 subtest
failures on the new rows, none of them passing by luck. Blocked before and now allowed or
blocked as they should be: 11 BLOCKED rows returned None ("# macOS" then ps -E, "# check ..."
then systemctl --user show-environment, "# run it" then systemd-run --pipe --wait printenv,
echo ${#PATH}; printenv, echo $#; printenv, gh api .../issues/1#c; cat "$PAPER_ENV_FILE",
echo a#b; printenv, two commands with a trailing comment, echo "$(date # )" whose printenv line
follows the comment, a comment holding quotes and a backquote, and a here-document heading
followed by printenv), plus 22 behind rtk proxy and 11 behind a keyring exec. 8 ALLOWED rows
returned a reason: echo ok # "$(printenv)", printf '%s' $'it\'s "$("printenv")"', ps -fu Eve
(two rows) and four commit-message patterns with a quoted here-document. 11 scanner-table rows
and 1 expected-pass-through row (a shell reading a quoted here-document in a substitution)
differed, and all 10 subtests of the five real commit messages (cf58426, 32a6cd2, a84fa7c,
88af2ba, b553826, verbatim, in the standard git commit -m "$(cat <<'EOF' ...)" and the gh pr
create --body forms) failed. The bug behind the first group is older than this work: tokenize()
joins lines with ";" and shlex reads a "#" anywhere, "$#" and "a#b" included, as a comment, so a
"#" dropped the whole rest of the command (found by the 2026-09-29 review).

Design: scan_shell() (substitution_bodies() now calls it) returns the bodies and the comment
spans from the same single pass. A comment is a "#" that starts a word (at the start of the text
or after a blank or one of ;&|()<>, and not right after a quote, escape or substitution), outside
quotes and outside a double-quoted substitution, to the end of the line; inside a $( ) body it
also runs to the end of its line, so a ")" in it closes nothing, and in a backquote body it ends
at the closing backquote. tokenize() removes the spans and sets shlex's commenters to "", and
command_segments() also reads the command as shlex did before (for commands up to 200,000
characters) and adds only the segments that reading lacks, so nothing the guard read before is
dropped. In a here-document body a comment hides the rest of that body, as before, not the text
after the terminator, so a script written through a here-document (its first line is #!) stays as
unread as it was. $'...' is data up to the first quote that no backslash escapes, except inside
double quotes. A here-document whose delimiter is quoted (also partly: <<'E'OF) is data: the body
that scan_shell returns loses it, keeping the operator line and the terminator; an unquoted
delimiter keeps its body, as at the top level. ps_shows_environment() skips the word after a
cluster that ends in a value option (ps -fu Eve). The three vacuous BLOCKED rows are marked as
regression rows and two are replaced by rows that pass at base, echo "$(echo \"$(printenv)\")"
and echo "$(base64 ~/.ssh/id_ed25519)".

Checked: oracle 65 cases, 0 mismatches; suites: only test_host_profile_copy_is_verbatim fails
(139 tests). Differential against base c26800f over 54,126 generated commands: 0 loosened, 0
unexplained reason changes, 0 newly blocked without a new form. Random grammar fuzz against
base, 160,000 strings in 4 seeds: 0 loosened, 0 exceptions. Real repository, by realistic shape:
785 fenced blocks, 4,338 fenced lines and 201 scripts written through a quoted here-document: 0
newly blocked; the synthetic shape of a whole script passed as one command string newly blocks 4
of 201, none of them a true positive and all long-standing readings that the old "#" had
hidden (three array literals x=(env -i ...) and a python set(...)). Commit messages: the
standard pattern newly refuses 0 of 842 (both guards refuse the same 8; the previous tip newly
refused 3 of the 842 and 6 of 1,865 across all refs); the top-level
git commit -F - shape newly blocks the two messages of this work that quote the new ps -E and
systemd-run rules in prose, by those rules and not by the comment change. 1 MB timing: a 840 KB
git commit -m "$(cat <<'EOF' ...)" takes 8.5 s at base and 8.9 s now (two readings would take
17.8 s, so commands above 200,000 characters are read once).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ocuments, systemctl show, quoted wrappers), correct the claims

Failed first (tests/test_secret_path_guard.py, run against the guard of 0ad4835): 154 subtest
failures on the new rows, each a reason expected and None returned. 41 BLOCKED and KEYRING rows,
plus 78 behind rtk proxy and 31 behind a keyring exec (numbers before the leading-redirection
rows, which failed the same way): a backquote in a single-quoted string
handed to bash -c, sh -c and eval; redirections between systemd launcher options
(systemd-run --user 2>/tmp/log --pipe printenv, ... > /tmp/log -E HF_TOKEN true); $(<.env),
$(< ~/.aws/credentials), x=$(<~/.netrc), echo `<.env`; unquoted here-document bodies that run
$(...) (cat <<EOF with value: "$(printenv)"); run0, systemd-inhibit and systemd-cat; a bare
systemctl show and show -p Environment; /usr/bin/sudo, /usr/bin/timeout, /usr/bin/nice; ps -CE
and ps -C -E; systemctl and systemd-run behind watch and flock inside a keyring exec (the wrapper
matrix rows). 3 scanner-table rows for here-document bodies failed too. The 19 new ALLOWED
controls (the alternatives the docs and the hint recommend, $(<version.txt), a quoted
here-document, path-qualified launchers of ordinary commands, ps -C python3, systemctl show
naming a unit) and the 12 new expected-pass-through rows passed before and after, as they must.

Design: scan_shell() also returns the backquotes that sit in single-quoted and ANSI-C strings
(tokenize() keeps them, so the shell that string is handed to sees them: item 9), the bodies of
an unquoted here-document scanned as double-quoted text with no closing quote (a "hd" frame, one
substring scan per body, nested at most 8 deep), and the quoted here-document bodies inside a
substitution that closes, which the tokenizer no longer reads in either of its readings (the
five real commit messages, and two commit messages of this work that the earlier guard refused
through shlex's quote parity, now pass in the standard pattern). wrapper_options() skips
redirections for the systemd launchers only: for timeout the duration in "timeout 5 > out cmd"
reads as a descriptor and the command was swallowed (the base differential caught it), so every
other wrapper keeps the walk it had. A token "(<" (shlex joins punctuation that touches) and a
segment that is only "< FILE" add a "cat FILE" segment, and a leading "< FILE cmd" adds "cmd < FILE",
in addition to what was read before.
The systemd launcher tables (run0 from parse_argv_sudo_mode v256-v262, systemd-inhibit,
systemd-cat) and the value options of later systemd-run releases cite the tag that first has
each (--root-directory v259, --output v261: the review found them absent from v256-v258).
systemctl_call() reads the verb and its arguments wherever the options stand, so a bare show
(or show -p Environment) with no unit is refused. prefix_end() compares the program name of a
wrapper, so path-qualified launchers count. ps_shows_environment() refuses an E word right
after -C. launched_commands() reads systemctl and the systemd launchers, and keyring_reason()
returns service_manager_environment for a started systemctl. The earlier reading of a command
that holds a # (or a protected backquote, or data) is kept beside the new one, for commands up
to 200,000 characters, so nothing the guard read before is dropped. The dated docs subsection is
rewritten with the corrected claims: the manager's "Started <unit> - <command line>" journal
line (measured on systemd 255.4, not _CMDLINE); --pipe returns the output, --wait shows terse
unit information and the output goes to the journal without either; the scope of "every rule
applies"; the alternatives considered for reading shell syntax (bashlex, tree-sitter-bash,
mvdan/sh, with the dates and licences read from PyPI and the GitHub API on 2026-09-29); the
residuals with their inert strings; and the systemctl hint now names the alternatives. Three
files that recommended the refused command (omniroute.service, token-report-refresh.service,
tools/token-report/README.md) keep it and add that it is for the operator's own terminal, with
the agent alternative.

Checked: oracle 65 cases, 0 mismatches; suites (4 named): only test_host_profile_copy_is_verbatim
fails (141 tests); the suites that read the three edited files (omniroute unit, token-report
units, ecosystem manifest, dashboard data, workflow hardening, adoption docs and status, 319
tests) pass. Differential against base c26800f over 59,530 generated commands: 0 loosened, 0
unexplained reason changes, 0 newly blocked without a new form. Random grammar fuzz against base,
240,000 strings in 6 seeds: 0 loosened, 0 exceptions. Mutation fuzz of strings the base blocks
(quotes, comments, backquotes, here-documents and launchers inserted), 8 seeds of 22,440
mutants: 1 differs (the run before the last change, seeds 21 to 28, had 3), every one a quoted
here-document whose body holds the printenv (bash runs nothing there). Walker comparison with base, 400,000 random cases: 0 mismatches. Real repository
by shape: 785 fenced blocks, 4,338 fenced lines and 201 scripts written through a quoted
here-document newly block 0; a whole script passed as one command string newly blocks 4 of 201
(array literals x=(env -i ...) and a python set(...), read as commands as before). Commit
messages: the standard pattern newly refuses 0 of 1,884 (all refs); two messages of this work that
the earlier guard refused there now pass; the unquoted-delimiter and the top-level git commit -F -
shapes read prose as commands as before (newly refused 8 and 4 of 843, by the # fix and by this
work's own new rules quoting themselves in prose). tests: check() runs without raising over 968
rows.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ommand takes data

The data rule of 172596e treated the body of a quoted here-document inside a double-quoted
command substitution as data whatever received the substitution. The coordinator's probes found
the cost: bash re-reads what a substitution prints as code for `eval "$(...)"`, `bash -c
"$(...)"` (also sh, zsh, behind env, timeout, nohup, xargs) and a code consumer nested in an
echo's substitution, and the base guard refused all of those as environment_dump while 172596e
allowed them.

Failed first (the new tests against 172596e): 196 failures and 2 errors, one failure being the
tolerated host-copy comparison. The 38 new BLOCKED rows failed directly (38), behind rtk proxy
(76) and behind a keyring exec (38); 27 consumer rows, 13 substitution-body rows and 2 oracle
groups failed; the two data-rule tests errored on the missing names.

Design: scan_shell() decides each quoted here-document only after the whole text is read.
data_consumers() reads the text once more with every returned substitution replaced by a marker
and the comments and top-level here-document bodies cut, finds the command each marker is an
argument of after the reserved words, assignments, wrappers and launchers the guard models
(receiving_command), and answers True only for git, gh, echo, printf, cat and tee. A shell, eval,
an interpreter, source, xargs, watch, ssh or any other program, an assignment, the command
position, a marker before the program word, a substitution nested in another, and any text the
guard cannot tell (markers in the text, more than 100,000 characters left after the cuts, more
than 16 launcher hops) keep the old reading: the body read as command lines and the tokenizer
input whole. Only a here-document at the command level of the returned body is data, so
`echo "$(echo "$(cat <<'EOF' ...)")"` keeps the strict reading. cat and tee count only for a
here-string word (`cat <<< "$(...)"`): a matrix of 39 heads x 18 launcher prefixes x 5
here-document forms x 8 payloads x 4 substitution forms (112,320 commands) showed the plain
allowlist loosening `cat "$(cat <<'EOF' ... a credential path ...)"`, which the base refused
through the reader rules, because a reader's operand is a file name; the operand forms now keep
the base verdicts.

Checked: kw/guard_probe_final.py base candidate reports loosened rows: 0 (172596e: 1); the
eight-consumer probe, with rtk proxy and keyring exec forms, 0 of 13 loosened (172596e: 13);
oracle 65 cases, 0 mismatches; the matrix loosens 0 rows for every head that is no data
consumer (172596e: up to 2,040 per head) and 0 for git, gh, echo and printf; mutation fuzz, 60
mutants per each of 397 blocked strings: 6 loosened at seed 7 and 2 at seed 11 (172596e: 387
and 378), each a mutant whose broken terminator or quote makes the rest of a git or echo
substitution here-document data that bash never runs, or a bash syntax error; differential
against c26800f over 59,592 commands from 936 rows: 0 loosened; the real-repository corpora
(fence lines and blocks, scripts through here-documents): 0 loosened; tests.test_secret_path_guard
29 tests, the host-copy comparison the only failure.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…and_too_large)

Claude Code's hooks documentation ("Timeouts", read 2026-09-29) says a timed-out command hook
does not block the tool call, and this hook's timeout is 10 s. Measured on this host on
2026-09-29 with one quoted word: 1,000,000 characters took 10.2 to 13.0 s in check() (base
c26800f and this version alike, over several runs), 600,000 took 3.8 to 4.4 s and 500,000 took
2.9 to 3.3 s, so a large enough command outlasts the hook and runs unread.

Failed first: test_the_size_limit_is_at_600000_characters and
test_a_command_over_the_size_limit_is_refused_without_echoing_it errored on the missing
MAX_COMMAND_CHARACTERS and the missing reason. The third test, a 500,000-character ordinary
command that must not be refused for its size, passed before and after, as it should.

Design: main() refuses len(command) > MAX_COMMAND_CHARACTERS (600,000 characters, not bytes)
before any rule reads the command, with exit 2 and one stderr line that names no command text:
"blocked (command_too_large)" and the hint to put the content in a file with the Write tool and
pass the path. check() is unchanged for size, so every table row behaves as before, and a
command of exactly 600,000 characters still reaches check().

Checked: a 700,000-character command is refused in under 5 s through the real hook subprocess
(exit 2, one line, neither the marker text nor its filler on stderr or stdout); 500,000
characters of ordinary shell (quoted words, arithmetic, redirections, separators) exit 0 with no
output inside 5 s and, with printenv appended, exit 2 with environment_dump; the boundary
test passes 600,000 characters to check() (mocked) and refuses 600,001 without calling it; 300,000
four-byte characters, 1.2 MB of UTF-8, pass. tests.test_secret_path_guard: 32 tests, the
host-copy comparison the only failure.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ries the new guard

docs/secret-storage.md, guard section, dated 2026-09-29 and no restructuring: one new item for
the rule that a quoted here-document is data only for git, gh, echo, printf, cat (as a
here-string word) and tee. It says why (bash runs what a substitution prints for a shell with
-c, eval, source, xargs or an interpreter), that every unlisted consumer keeps the stricter
reading so prose in a quoted here-document inside them can still be refused, what the guard
still does not read (an allowed consumer's output used later, a shell reading the quoted
here-document itself), and the measured matrix. The here-document item now says "where the
command takes data", and the timeout item gains one sentence for command_too_large with the
measured timings (1,000,000 characters of one quoted word: 10.2 to 13.0 s; 600,000: 3.8 to
4.4 s; 500,000: 2.9 to 3.3 s), replacing the "not fixed" note that described the gap. The
guard's own docstring says the same in two sentences.

adoption/hooks/claude/SHA256SUMS: only the guard's line changed, to
914342a040d9632b9aedc435e58617a78808b5f5f6fcbe88a9460edda41be75c (91,845 bytes); sha256sum -c
over the file reports every line OK.

Checked: kw/guard_oracle.py exits 0 (65 cases, 0 mismatches); kw/guard_probe_final.py reports
loosened rows: 0 with the timing rows of the last round unchanged (12,000 here-documents 0.19 s,
20,000-deep quoting 0.66 s, 60,000 systemd-run 0.58 s); tests.test_secret_path_guard,
test_credential_status, test_credential_tools and test_install_claude_profile ran 148 tests, the
host-copy comparison the only failure; 305 further tests (docs consistency, adoption status,
blind checkout, effort guard, release pins, landscape) pass; scripts/validate.py lists only the
manifest hash and byte mismatches of the pinned files this branch changed (SHA256SUMS, the
guard, its tests, docs/secret-storage.md, and the two files 172596e changed without a re-pin,
omniroute.service and tools/token-report/README.md).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…e inventory entry)

A later inventory entry, canary-e2e (a disposable synthetic proof key: class test_canary, status
test_only), declares the variable CANARY_E2E_KEY, and test_inventory_secret_names_are_all_guarded
requires every inventory name to be in SECRET_NAMES, so it would fail there without the name.
It is added now, ahead of that entry (which belongs to another branch and is not touched here),
and is refused like every other name.

Failed first (the two new rows against 914342a0, the guard of 82767ca): 9 failures, one being the
tolerated host-copy comparison; the rows failed directly (2), behind rtk proxy (4) and behind a
keyring exec (2). The four forms `echo "$CANARY_E2E_KEY"`, `echo $CANARY_E2E_KEY`, the
os.environ lookup and `rg -n CANARY_E2E_KEY` all returned None; they now return
secret_variable_reference (three) and secret_name_search, and `rg -n MY_CANARY_E2E_KEY_HINT`
still passes (the word boundary).

Design: one entry in SECRET_NAMES with a comment; the expansion, lookup and search rules and
test_every_secret_name_is_caught_by_a_search pick it up from the tuple. Explicit BLOCKED rows for
`echo "$CANARY_E2E_KEY"` and `rg -n CANARY_E2E_KEY`, so the wrapper tests also run them behind rtk
proxy and a keyring exec. No docs sentence lists where the names come from (only the comment above
the tuple does), so the docs are unchanged. The inventory tie holds with and without the later
entry's variable (29 inventory names, 30 unique guard names, none missing). SHA256SUMS carries the
final guard hash for this and the earlier changes of this round.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The independent verification review asked for a limit low enough that the tokenizer's two readings
of a command (with and without shlex's comments) always both run, so that no reading is skipped
above some length. Measured on this host on 2026-09-29, one quoted word costs 0.6 s at 200,000
characters (about 1.5 s for both readings), 2.9 to 3.3 s at 500,000, 3.8 to 4.4 s at 600,000 and
10.2 to 13.0 s at 1,000,000 (base c26800f and this version alike), against the hook's 10 s.

Failed first: the size test that pins the limit failed (600000 != 200000) once it asked for
200,000; the others use the constant and follow it. Tests: a 250,000-character command is
refused through the real hook (exit 2, one line, no command text) inside 5 s; a 150,000-character
ordinary command exits 0 with no output inside 5 s and, with printenv appended, exits 2 with
environment_dump; the boundary test passes 200,000 characters to check() (mocked) and refuses
200,001 without calling it; 150,000 four-byte characters (600 KB of UTF-8) pass.

The constant's comment, the guard docstring and the docs sentence carry the new limit and the
timings; SHA256SUMS carries the guard's new hash (a stale one fails four tests of the other
suites).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…acktracking regular expression

The independent verification review found that `'ps ' + 'E' * 70000 + 'q; printenv'` took 13.2 s in
check() against 0.08 s at the base guard: PS_BSD_CLUSTER, `^[aAcfhjlmrsStTuvwxXLnE]*[eE][aAcefhjlmrsStTuvwxXLnE]*$`,
has two overlapping quantifiers that fail in quadratic time, and a hook that outlasts its 10 s
timeout does not block the call. Reproduced here at 15.2 s.

Failed first: the new timing row "ps cluster of 70,000 E and a letter that is no flag" failed at
15.2 s against the 3 s bound, and the equivalence test errored on the missing function.

Design: is_ps_bsd_cluster(word) is PS_BSD_LETTERS.issuperset(word) and an `e` or `E` in it, the
same language as the regular expression (checked on every word of up to 5 letters over nine
letters, 66,430 words, and on 13 words the guard meets). Five timing rows for ps clusters of
70,000 characters (a lead of E, of e, of a, a trailing E, a letter that is no flag) join the
PATHOLOGICAL table with their verdicts and the 3 s bound, and a test times every compiled pattern
of the guard (module level, the scan and store tables: 52) on 70,000 repeats of 12 characters
with a lead and a tail that make a match fail late, each under a quarter of a second. Negative
control: the same test against the guard before this change fails naming PS_BSD_CLUSTER at 14.6 s.
The audit script that found nothing else (PAREN_INPUT is slow under search, 6.4 s, but the guard
only calls it as fullmatch, which is linear) also timed 2,132 check() shapes, a first word
(ps, env, sudo, systemd-run, systemctl, keyctl, rtk, grep, git, bash -c ...) followed by 70,000
repeats of a character or a token: the slowest were the two ps shapes, everything else 0.3 s.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…the only exemption

The independent verification review of 172596e (GPT-6, effort max, read-only) found six blocking
block-to-allow regressions, five in the general here-document machinery: a quoted delimiter made
the body of a here-document inside a double-quoted substitution data, and each imperfection of
the bash-exact parsing that needs erased executable text. Reproduced here as inert strings, each
an environment_dump at the base guard c26800f and None at 172596e: `echo "$(bash <<'EOF'` with
`echo "$(printenv)"` in its body (a shell reads the body); a `<<$'EOF'` delimiter; a body line
`text\` that the guard's backslash-newline join glued to its terminator; an arithmetic command
`((1 << "2"))` taken for a `<<`; a quoted `#` (`printf ' #x'`) taken for a comment that hid the
rest of a here-document body. The design change of the coordinator replaces the general reading.

Failed first (the new tests against 3713250): 125 failures and 129 errors, one failure being the
tolerated host-copy comparison: the recognizer table (61 rows), the consumer table (60 rows), the
41 new BLOCKED rows directly (29), behind rtk proxy (58) and behind a keyring exec (29), the
plain-scan rows of the bodies table (6), the ANSI-C rows (7), a known-bypass row, the neutral
reading and the joined-line test.

Design. scan_shell() no longer reads here-documents at all: the delimiter regular expressions,
heredoc_delimiter, heredoc_end, the terminator index in the scanner, the here-document data
spans, the comment-in-a-body rule and the `hd` frame are deleted, so a here-document's lines are
command lines as the tokenizer has always read them at the top level (a `#` hides its own line,
a `"$(...)"` in a body line is a substitution). The one exemption is idiom_spans(): a
double-quoted word that is exactly `"$(` blanks `cat` blanks `<<` [`-`] blanks `'IDENT'` blanks
newline, the body, the FIRST line that is exactly IDENT (for `<<-` after its leading tabs, no
other trimming, no carriage return), blanks and newlines, `)"`, standing alone as a word. It is
looked for in the text as written, before check() joins backslash-newline pairs, with a line
index built once and a tail memoised per terminator, so no head costs a rescan. exempt_idioms()
then reads the text once more with each idiom replaced by a marker and asks the command that
receives it: git, gh, echo or printf after the reserved words, assignments, wrappers and
launchers the guard models, with the idiom a word of its own after the program, not inside the
body of a substitution, and the text readable by shlex without its fallback. Then neutral_reading()
puts a neutral word in its place for the COMMAND reading only; the raw-text rules (store paths,
secret names and expansions, /proc) read the whole text as before. Everything else, a shell,
eval, an interpreter, source, xargs, watch, ssh, cat, tee, an assignment, the command position,
`\EOF`, `"EOF"`, `$'EOF'` and unquoted delimiters, text on the operator line, text after the
terminator, an unterminated body and a word glued to `--body=`, keeps the old reading, the body
lines read as commands. The legacy-reading cutoff is deleted with the machinery. shlex knows no
ANSI-C quoting, so tokenize() keeps each `$'...'` string as one word (harmless
`printf '%s' $'it\'s #\nprintenv\n'` passes; `bash -c $'printenv'` is read as `bash -c printenv`).

Checked at this state: the oracle 65 cases, 0 mismatches; kw/guard_probe_final.py base candidate
`loosened rows: 0`; the eight consumer variants and their rtk proxy and keyring exec forms 0 of 13
loosened; the reviewer's five here-document strings and the coordinator's eight consumer forms
are BLOCKED rows and environment_dump; the four suites 151 tests, the host-copy comparison the
only failure; 14 mutation controls of the recognizer, the consumer decision and the tokenizer
change each make a test fail (two survivors found by the first run, a plain `<<` terminator with
a leading tab and the second tokenizer reading, got a row each). The launcher-redirection strings
and the corrections of the review's other findings follow in their own commits.

Residual gaps, recorded as EXPECTED_PASS_THROUGH rows: a git or gh option that runs its value
(`git rebase --exec "$(cat <<'EOF' ...)"`, `gh alias set`), an echo piped into a shell or written
into a script that runs later, and an unquoted here-document that expands `$(...)` between single
quotes (the base guard passed it too). Friction, in the strict direction: prose in a body that
starts a line with `printenv` is refused outside the idiom, and so are `--body="$(...)"`, `cat`
and `tee` as consumers and a text with an unbalanced apostrophe before the idiom.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The review of 172596e found that `env -u < "$PAPER_ENV_FILE" UNUSED cat` and `systemd-run --pipe
--unit < "$PAPER_ENV_FILE" demo cat` are credential_file_read at the base guard and None at the
tip: the index-based launcher walk took the redirection operator for the value of `-u` and of
`--unit` and dropped the segment that held the credential-file operand. Reproduced with a
matrix of redirection positions: for env, env -i, systemd-run, sudo, timeout, timeout -s, nice,
rtk proxy, a keyring exec and four chains (sudo env, env sudo, nohup timeout), every position
between the hops, each of `<`, `<<<`, `2>`, `>`, `>>` and `&>`, to a credential pointer, `.env`
and `~/.aws/credentials`, 684 combinations: 8 rows looser than the base guard, all `<` or `<<<`
to the pointer in an option-value slot (`env -u < ...`, `systemd-run ... --unit < ...`, the same
behind sudo and env).

Failed first (the new tests against the guard of the previous commit): 170 failures, one being
the tolerated host-copy comparison: 152 subtests of the position matrix, 13 reviewer strings, and
the wrapper tests behind rtk proxy (2) and a keyring exec (1) and the direct table (1).

Design, following the coordinator's instruction: the raw segment is read as written beside every
launcher walk (command_segments appends it before any walk), so a redirection is checked
wherever a walk would put it; and no option-value slot swallows a redirection any more.
skip_redirections() steps over a redirection operator and its target, and wrapper_options,
env_command_start, rtk_command_start and prefix_end use it wherever they look for an option, a
value or the started command (only an operator token starts one: a bare number may be a value,
`nice -n 5 > out cmd`). An input redirection that was stepped over is appended to the command that
gets it, as if it stood after it (launcher_chain now returns it), so `env -u UNUSED < .env cat` is
read as `cat < .env`, and a keyring exec and a leading here-string do the same. Nothing is
carried for a number before an operator: `sudo 2>/dev/null -u root printenv` stays a recorded gap.

Checked at this state: for every launcher, position and operator the verdict equals the verdict of
the same command with the redirection written last, 684 combinations, 0 mismatches; against the
base guard 0 rows are looser and 166 are stricter, each the operand of a `cat` that the base
guard only saw when the redirection stood after the command; kw/guard_probe_final.py and the
oracle unchanged (`loosened rows: 0`, 65 cases, 0 mismatches); the smoke row of each reviewer
string is credential_file_read; the four suites 153 tests, the host-copy comparison the only
failure; 22 mutation controls (the earlier 14 and eight for the walkers, the raw segment and the
carried redirections) each make a test fail. `nice > /tmp/out -n 5 cat .env`, a recorded gap of
172596e, is now a BLOCKED row. The redirection rows also cover a launcher with no command
(`sudo < "$PAPER_ENV_FILE"`), which only the raw segment sees.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…oosenings and the ps friction

The review's other findings, each a correction of a claim, not of a rule.

systemd-run provenance. The guard's hint, docstring, comments and docs said that the manager logs
systemd-run's command line, an `-E NAME=value` included, in the journal. Read from upstream
(systemd v255, src/run/run.c, fetched 2026-09-29): the default Description, which the manager
logs as `Started <unit> - <description>`, is `quote_command_line(arg_cmdline)` (lines 1940 to
1951), and arg_cmdline is the command and its arguments after the options; `-E` fills
arg_environment (line 348), which is appended to the start message as the unit's Environment
property (lines 853 to 866). So a value given with `-E`, `--setenv` or `-p Environment=` is in
the command line of the systemd-run process, which the process listing shows while it runs, and
in the transient unit's Environment property on the user bus (`systemctl --user show -p
Environment UNIT`, and any bus client); it is not in the journal, while a value written after the
command is (a recorded gap). The rule, a secret name on a systemd-run line is refused, is
unchanged. The verbatim commit message of cf58426 that the tests carry as a real message keeps
its old wording, as it is a historical text.

Failed first: test_the_systemd_run_hint_says_where_a_value_on_the_command_line_goes failed against
the previous guard (the hint named the journal); the rest of this commit is documentation and
rows that the previous guard already satisfied (ALLOWED rows for `ps -fu steve`, `ps -fu eve` and
`ps -u steve` pass there too).

The list of loosenings. The base guard c26800f refuses `ps -fu steve` and `ps -fu eve`, because
it reads the user name as a dashless BSD cluster with an `e` (`ps -u steve` it passes); `-u`
takes a value, and this work lets them pass. Together with the canonical idiom of the
here-document item, that is the only loosening against the base guard: docs/secret-storage.md now
says so in one item and the ALLOWED rows carry it. `ps -CEmacs` stays refused, documented as
friction (on procps `-C` takes a command name: `ps -C emacs`).

Checked: the oracle 65 cases, 0 mismatches; kw/guard_probe_final.py `loosened rows: 0`; the four
suites 154 tests, the host-copy comparison the only failure.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ected measurements

My own adversarial probe of the exemption before the final pass, nested and wrapping constructs
against the base guard, found that `eval $(echo "$(cat <<'EOF'` + `printenv` + `EOF` + `)")` passed
(the base guard passed it too, so it is no regression, but bash runs `printenv` there): the
tokenizer splits an unquoted `$(echo IDENT)` into a segment of its own whose command is echo, and
exempt_idioms() judged the idiom by that segment, while the substitution's output goes to eval.
The same held for `bash -c $(echo IDIOM)`, `bash <<< $(echo IDIOM)`, `x=$(echo IDIOM); eval $x`,
`xargs sh -c $(...)`, `python3 -c $(...)`, `eval $(git log -1 --format=%B IDIOM)`, a subshell or
brace group around any of them, and `echo $(git commit -m IDIOM)`. The dq-substitution rule of the
previous commit covered only the substitutions inside double quotes.

Failed first (the new tests against the guard of the previous commit): 68 failures and 16 errors,
one failure the tolerated host-copy comparison: 15 rows of the exemption table, the 14 new
BLOCKED rows directly (13), behind rtk proxy (26) and behind a keyring exec (13), and the two
tests of the new scan result (10 and 6 errors).

Design: scan_shell() now returns a fifth result, the (start, end) of the text inside each
outermost command substitution, `$(...)` or a backquote pair, quoted or not (a counter of open
substitution frames, kept right through the `$((printenv) )` conversion and an unterminated
substitution); exempt_idioms() refuses any idiom whose marker lies inside one, which replaces the
bodies-only rule. A subshell `( ... )` and a brace group are no substitution: `( cd sub && git
commit -m IDIOM )` stays exempt. The probe's 28 rows: 0 looser than the base guard; every nested
form above is now an environment_dump; `git commit -m IDIOM` and `echo IDIOM` pass; the pipe into a
shell stays a recorded gap.

Docs: the exemption item names the rule (`eval $(echo "$(cat <<'EOF' ... EOF)")`), and the measured
numbers of the item are corrected to the final ones: the consumer matrix is 40 consumers x 18
launcher prefixes x 5 here-document forms x 8 payloads x 4 substitution forms = 115,200 commands
(the item said 39 and 112,320), the commit history is 1,956 messages (13 refused by the base guard,
10 by the guard, all by raw-text rules, 4 that the base refused pass), the differential over 60,493
derived commands loosens 102 rows and every one is a `ps -fu steve` or `ps -fu eve`, and the
mutation fuzz of 24,840 mutants of 414 blocked strings loosens 38 to 44 per seed, of which 36 to 41
are variants of those two rows and 2 or 3 are mutants whose broken terminator leaves a real quoted
body (the item said 24,720 mutants, 412 strings and 2 or 3, without the ps rows).

Checked: the oracle 65 cases, 0 mismatches; kw/guard_probe_final.py `loosened rows: 0`; the four
suites 155 tests, the host-copy comparison the only failure; 22 mutation controls each make a test
fail (the rule that this commit strengthens is killed by the new rows).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…her strings are table rows

Housekeeping the design change left: the guard's module docstring still described the superseded
rule ("a quoted here-document inside a substitution whose command is git, gh, echo, printf, cat or
tee"), and said backslash-newline is joined first without saying that the canonical idiom is looked
for before that. It now says: a here-document's lines are command lines like any other; the one
exemption is a strict canonical idiom, a double-quoted word of exactly the shape
`"$(cat <<'EOF'` newline, body, `EOF` newline, `)"`, behind git, gh, echo or printf, whose body is
data for the command reading while the raw-text rules still read it. Three test comments that
named "data consumers" or "the first repair round" follow.

The two launcher strings of the review, `env -u < "$PAPER_ENV_FILE" UNUSED cat` and `systemd-run
--pipe --unit < "$PAPER_ENV_FILE" demo cat`, were checked by a test that asserts only that they
are refused; the coordinator asked for them in the tables, so they are BLOCKED rows
(credential_file_read, as at the base guard), where the wrapper tests also run them behind rtk
proxy and a keyring exec.

No behaviour changes: the four suites and the oracle pass as before, and the guard's hash in
SHA256SUMS moves with the docstring.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A grammar fuzz of my own, launcher chains with redirections at random places and idioms in random
placements, generated against the base guard (10 seeds of 60,000 commands: 5 with the launcher
walkers of the previous commits), found 4 commands in 120,000 that the base guard blocks and the
guard allowed: a keyring exec with a redirection between the variable name and the `--`, for
example `python3 scripts/kernel_keyring.py exec tavily_api_key TAVILY_API_KEY < x.txt -- set`
(base: environment_dump_in_keyring_exec). The base guard dropped that redirection when it took the
command after `--`; the walkers of this work carry a skipped input redirection to the command it
belongs to, and is_environment_dump() counted the carried `< x.txt` as arguments of `set`
(`set a b` sets positional parameters), so the dump was no dump. `set < FILE` and `set > FILE`
still print every variable.

Failed first (the new rows against the guard of the previous commit): 53 failures, one the
tolerated host-copy comparison: 16 BLOCKED rows directly, 28 behind rtk proxy, 8 behind a keyring
exec. Four of the new rows contain a keyring exec of their own and live in KEYRING_BLOCKED, which the
wrapper tests do not prefix with another one.

Design: the module docstring already says "redirection operands are never taken for arguments";
is_environment_dump() now reads set, export, declare and typeset through command_arguments(), which
drops the redirections, so `set > /tmp/vars`, `set < /dev/null`, `export -p > FILE`, `declare -p >
FILE`, `typeset -p 2> err` and the keyring forms are an environment_dump, while `set -e >
/dev/null`, `export FOO=1 > /dev/null` and `declare -a items > /dev/null` (a real argument) pass.
This only tightens: the base guard passed the redirected forms.

Checked: 10 seeds of the grammar fuzz, 600,000 commands, 0 looser than the base guard (seeds 1 and 2,
which found the four, included); the four suites and the oracle unchanged; the docs sentence is in
the launcher-redirection item.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The here-document item of the guard section still quoted the numbers of an earlier commit of this
series (1,956 commit messages, 13 and 10 refused, 4 passing, a differential over 60,493 commands, mutation
fuzz of 24,840 mutants of 414 strings). Replaced with the numbers measured at 660e681, the tip whose
guard is byte-identical to this one: 1,960 messages, 10 refused (base 14), 5 that the base guard refused
now pass, a differential over 62,365 derived commands that loosens 102 (all `ps -fu steve` or
`ps -fu eve`), a grammar fuzz of 10 seeds of 60,000 launcher chains that loosens none, and mutation fuzz
of 25,140 mutants of 419 blocked strings that loosens 33 to 44 per seed (29 to 38 `ps` variants, 4 to 6
mutants whose broken terminator leaves a real quoted body that bash prints and does not run). The
loosenings item says 5 messages at 660e681 instead of 4. Two over-long lines of the item re-wrapped.

Docs only: the guard is unchanged (sha256 72d43a2a15c4ccc002a13d9747ed818ee51832084811753912d4e9aa4ebf347e,
94,067 bytes), so the SHA256SUMS line stays. The four suites: 155 tests, one failure, the tolerated
host-copy comparison (the host profile copy is not reinstalled), 3 skipped.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…'s history

The measurement of the here-document item said "1,960 commit messages reachable there" at 660e681.
The script that counts them (git log --all) reads every ref of the shared repository, including other
branches and worktrees, so the number is the distinct messages on all refs when measured (1,964 on
the run that found this) and grows without any change to the guard. The item now says that, names the
guard as the one of 660e681, and the loosenings item points at that measurement instead of repeating
"at 660e681". Re-measured on the current refs: 10 messages refused by the guard, 14 by the base guard,
5 that the base guard refused now pass (3899d49, 2f097bf, 172596e, c3bdbaf, 5b96461), so the
sentences hold.

The timeout item claimed that both readings of a command always run inside the hook's 10 s at the
200,000-character limit. It now carries the measurement behind that: seven inputs of about 199,000
characters that make every reading run (double-quoted text with a `#` throughout plus one idiom or
1,000 idioms, the same single-quoted, ANSI-C strings, one quoted word, many backquotes, many subshells)
took 0.2 to 3.3 s in check() on this host, the slowest being the double-quoted text plus one idiom.
Three over-long lines (the loosenings item and two in the timeout item) re-wrapped.

Docs only: the guard is unchanged (sha256 72d43a2a15c4ccc002a13d9747ed818ee51832084811753912d4e9aa4ebf347e,
94,067 bytes) and the SHA256SUMS line stays.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…read as command lines behind every consumer

The second verification review (GPT-6 through the gateway lane, of 50ca6ca) found six more blocking
regressions, five of them inside the strict canonical-idiom exemption that replaced the general
here-document reading: an idiom head inside another here-document's body that swallowed its
terminator, an idiom inside an unquoted here-document that a shell reads, an echo piped into a shell,
`git rebase --exec` (a git option that runs its value), a case pattern's `)` that closed a
substitution early, and a process substitution that runs a shell. Each is an environment_dump at the
base guard and None at that tip. The coordinator's decision: the exemption is a net risk for its
friction benefit, so it goes, and the guard is tightening-only again. Two reviews and twelve
block-to-allow regressions show that telling a data consumer from a code consumer by the words of one
command line is what fails.

Failed first (the six reviewer strings as BLOCKED rows against 50ca6ca): 24 failures, 6 direct, 12
behind rtk proxy and 6 behind a keyring exec; the two `ps -Ccat e` strings of the same review are the
next commit.

Design: idiom_spans, exempt_idioms, receiving_command, neutral_reading and everything only they used
(IDIOM_HEAD, IDIOM_TAIL, IDIOM_FOLLOWERS, IDIOM_CONSUMERS, RESERVED_STARTERS, SUBSTITUTION_MARKER,
MAX_LAUNCH_HOPS, the bisect import, and scan_shell's fifth result with its open-substitution counter)
are deleted. check() reads the joined command as it did before the idiom existed: a here-document's
lines are command lines, quoted delimiter or not, behind git, gh, echo and printf as behind a shell,
eval, an interpreter, source, xargs, watch and ssh. The raw-text rules, the double-quoted
substitution reading, comments, the launcher redirection handling, the systemd-run family, ps -E,
systemctl and CANARY_E2E_KEY stay.

Tests: the six strings and 16 documented-friction rows are BLOCKED rows (a commit message or pull-request
body whose prose has a line that reads as a dump; a literal `$(printenv)` in a quoted here-document
in file-writing, Python, commit-message and PR-comment workflows; Python's `set()` after a comment
line in an interpreter's here-document), ALLOWED has the workarounds (`git commit -F`, `gh ... --body-file`)
and a prose message that passes; the five idiom tests and their tables are gone, the real commit
messages of this repository are recorded as the friction they now are (each refused with its reason),
and a matrix of 26 consumers, 7 delimiter forms and 10 dumps (1,820 inert strings) must all be refused,
so an exemption that returns for one consumer fails there.

Checked: oracle phase base 65/65; the probe with the base guard `loosened rows: 0`; the four suites 151
tests, the one failure the tolerated host-copy comparison; every reviewer string of both reviews is
blocked. Of 2,015 distinct commit messages on all refs, in the `git commit -m` and `gh pr create --body`
patterns, the guard refuses 36 and the base guard 16, and none that the base guard refuses passes. The
seven worst-case shapes of about 199,000 characters now take at most 1.0 s (3.3 s with the third
reading). The guard section of docs/secret-storage.md states the friction and the workaround, and
its loosenings item says one; SHA256SUMS carries the new guard hash.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…two readings and either one refuses

The second verification review of 50ca6ca found that removing `-C` from the value-taking options of a
cluster lost the procps reading: `ps -Ccat e` and `ps -fCcat e` (environment_dump at the base guard)
returned None. On procps `-C cmdlist` takes a command name (`cat`) and the BSD `e` after it shows the
environment; on macOS `-C` is a flag, so `-Ccat` is `-C -c -a -t` and `t` takes `e` as a tty. The
coordinator's instruction: keep both readings and refuse when either shows the environment.

Failed first (the new rows against 752def7f): 16 failures, 4 direct (`ps -Ccat e`, `-fCcat e`,
`-Ccat eww`, `-Ccat E`), 8 behind rtk proxy and 4 behind a keyring exec. `ps -Cnginx e` and `-fCnginx
auxe` were refused before by chance (the `g` of `nginx` is a value letter) and stay refused.

Design: ps_reading_shows_environment(words, procps) is one host's reading, ps_shows_environment() the
union. procps (PS_ARG_OPTIONS, `-C` a value option, dashless BSD words anywhere, no dashed `-E`) and
macOS (PS_MACOS_ARG_OPTIONS, `-C` a flag, dashed `-E`, a dashless option string only as the first
argument) each keep their own state of which word a value option takes. The `-C -E` special case of the
hybrid reading is gone: the macOS reading finds it. Sources: procps-ng 4.0.4 ps(1) on this host (`-C
cmdlist`, `e` "Show the environment after the command", no `-E`), and Apple adv_cmds ps/ps.c at
60bc9ebf (fetched 2026-09-29): PS_ARGS `aACcdeEfg:G:hjLlMmO:o:p:rSTt:U:u:vwx`, `case 'C': rawcpu = 1`,
`kludge_oldps_options(..., argv[1], ...)` applied to argv[1] only, and a later non-option word is a
process id or an "illegal argument". Checked on this host: `ps -Cbash u` honours the BSD `u` after
the glued command name and `ps -fC bash u` reports conflicting format options, so a dashless word
after a UNIX option is a BSD option on procps. `ps -CEmacs` stays refused (friction: the macOS
reading; write `ps -C emacs`).

Loosening, the one this work makes, now stated generally: a value after a clustered value-taking
option is no BSD flag cluster (`ps -fu steve`, `-fo user`, `-ft e`, `-fU steve`, `-fC e`); the base
guard skipped the value after a stand-alone `-u` but not after `-fu`. A differential over 478,915 ps
command lines (every one of up to three words from a vocabulary of 65, plus 200,000 random ones of
four to six words) loosens 6,673 against the base guard, every one holding such a cluster with its
value (0 outside that family); the reference two-host model of the tests (a per-letter state machine)
and the guard agree on all 478,915. The tests carry that reference over 160,434 lines (a vocabulary
of 54), ALLOWED rows for the value forms and BLOCKED rows for the union.

Checked: oracle 65/65; probe with the base guard `loosened rows: 0`; the four suites 152 tests, the one
failure the tolerated host-copy comparison. The docs replace the ps item and the loosenings item, and
record macOS's legacy mode (`-e` read as `-E` when ps.c's `u03` is off) as a gap. SHA256SUMS carries the
new guard hash.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Scout and others added 12 commits October 1, 2026 01:47
… refused as command_too_complex

The second verification review found two commands far below the 200,000-character cap that outlast the
hook's 10 s timeout, where a timed-out PreToolUse hook blocks nothing: a keyring exec with 1,200
`python3` words (9,645 characters; the analysis read every suffix of them as a command, 5.8 million
characters through shlex: base 2.1 to 2.8 s, 50ca6ca 12 to 13 s) and 195,000 emoji in five nested
double-quoted substitutions (195,068 characters; four-byte characters cost four times as much to
tokenize and the text was read at every level: base 1.6 s, 50ca6ca 15.7 to 17.9 s here, over 25 s
in the review). My own scan found the same class elsewhere: 4,000 distinct keyring execs (25 s), 3,500
keyring execs of one variable, each starting a distinct program (268 s: a quadratic test of mentions
against spans and a scan of the whole text for each started command), a chain of 6,000 nested keyring
execs (a copy of the rest of the words at every hop, past 30 s), and a keyring exec whose started
command is a 180,000-character argument (6.4 s).

Failed first (the new tests against d46f6c0c): the six unit tests 2 failures and 10 errors (no
WorkBudgetExceeded, WORK_LIMITS, storage_width or start_work, expand kept 3 copies of one segment,
no command_too_complex hint); the timing rows in the tests' own child process: 1,200 words 12.48 s,
600 words 3.16 s, 5,000 words past the 45 s limit, the emoji input 17.20 s, 120,000 euro signs 11.0 s.

Design: check() starts a budget, reads the command, stops it (start_work, stop_work; outside a check
call nothing is counted). WORK_LIMITS: (a) characters passed to shlex over every reading and nesting
level, each at its storage width (1 byte ASCII and Latin-1, 2 for the rest of the BMP, 4 beyond: shlex
costs 0.41, 0.85 and 1.70 s on 200,000 characters of one quoted word), 400,000; (b) texts read 10,000,
words emitted (each segment its words plus one) 1,000,000, launched-command reads of the keyring
analysis 500. lex() charges before shlex runs, command_segments() charges each text and each emitted
segment, launched_commands() and keyring_reason() charge each read. Over a limit check() raises
WorkBudgetExceeded (args[0] names the counter), never a reason, so no caller can take an unread command
for an allowed one; main() prints `blocked (command_too_complex)` with one line, no command text, and
the hint to split it or use the Write tool, and exits 2. Algorithmic fixes that need no budget:
expand() and the keyring analysis drop identical segments, keyring_reason() scans the text once for
each pattern (a cache), and mentions_injected_variable() tests a mention against the starts of the
exec arguments (a set) instead of every span. The dead `strict` parameter of lex() and tokenize() went
with the idiom code.

Limits are from measurement, not a guess: each counter alone holds its worst case near a second
(398,000 characters 1.1 s, 9,900 texts 0.3 s, 1,000,000 words 0.2 s, 490 reads 0.03 s); the largest real
command of this repository (an 82,000-character script through a here-document) spends 35% of the
characters and 1% of the words, the largest commit message 57%, none of 785 fenced blocks, 4,338
fenced lines, 201 scripts or 1,877 commit messages comes near a limit; 500 random mixes of 22
adversarial building blocks at 199,000 characters took at most 0.9 s, and 64 shape families at 25,000
to 199,000 characters at most 1.1 s. The reviewer's two inputs are refused in 0.15 s and 0.04 s, their
600- and 5,000-word variants in 0.15 s.

Tests: 33 timing rows with the 3 s bound (the reviewer's two, 600, 60 and 5,000 words, euro signs,
nested and distinct keyring execs, a chain of 6,000, the launched programs and the 100,000-character
argument, each with the deterministic verdict, `command_too_complex` for what the budget refuses),
counters tripped one at a time and the same shapes just inside, the storage width, the accounting
of one expand() call, the deduplication, main() and the real hook on both inputs (under 3 s), and
check() may raise only WorkBudgetExceeded on a pathological row. Rows of the earlier rounds that the
budget now refuses were resized (a 160,000-character 20,000-deep nesting is refused, 5,000 deep
reports its dump).

Checked: oracle 65/65; probe with the base guard `loosened rows: 0` (its 1 MB quoted word now
refused in 0.15 s, was 10.4 s); the four suites 157 tests, the one failure the tolerated host-copy
comparison. The guard section states the counters, the numbers and the friction; SHA256SUMS
carries the new guard hash.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… from `-E NAME`

The second verification review's note: the provenance still conflated an explicitly supplied value with
value-less environment forwarding, in docs/secret-storage.md and in the module (the hint and the
docstrings). Read from systemd v255 (fetched 2026-09-29): `-E NAME=value` puts the value in the
systemd-run process's own argv (the process listing shows it while systemd-run runs) and in the transient
unit's Environment property; `-E NAME` puts only the name in argv, and systemd-run takes the caller's own
value from its environment (`strv_env_replace_strdup_passthrough`, src/basic/env-util.c:417, called for
`case 'E'`, src/run/run.c:348), so that value reaches the Environment property and no argv. The property
is appended to the start message over the user bus (`arg_environment`, run.c:853-866), which
`systemctl --user show -p Environment UNIT` and any bus client read. Not the journal: the unit's
description, which the manager logs as `Started <unit> - <description>`, is the started command and its
arguments after the options (`quote_command_line(arg_cmdline)`, run.c:1940-1951), so neither `-E` form is
in it; a value written after the command is (a recorded gap, unchanged).

Failed first: test_the_systemd_run_hint_says_where_a_value_on_the_command_line_goes against the
previous hint (which named the process listing and the user bus for every form): 1 failure.

Changed: the hint of `secret_variable_on_command_line` (it now says `-E NAME=value` puts the value in the
command line, `-E NAME` forwards the caller's value, and either way it lands in the Environment property
on the user bus, readable with `systemctl --user show -p Environment UNIT`; still no word about the
journal), the docstring of systemd_run_sets_secret, the module docstring and the docs paragraph of the
systemd launchers, with the three source lines cited. The rule itself is unchanged: the same names on
the same forms are refused (the review's `systemd-run --user -E GH_TOKEN true` is a row of the hint test).

Checked: oracle 65/65; probe with the base guard `loosened rows: 0`; the four suites 157 tests, the one
failure the tolerated host-copy comparison. SHA256SUMS carries the new guard hash.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…gacy `-e` as a recorded gap, a mutation the word limit survived

Tests only; the guard is unchanged (sha256 678cd174f3c8bfbe936a8485e21902ed82c2580fd5cc21c4a6b4bbf0a21a9b29,
93,563 bytes), so the SHA256SUMS line stays.

- BLOCKED: `ps` followed by 70,000 `E`, a letter that is no flag and a dump, the first review's input in
  full (the timing row PATHOLOGICAL already bounded its time; the BLOCKED table also runs it behind rtk
  proxy and a keyring exec).
- EXPECTED_PASS_THROUGH: `ps -e`. macOS's ps.c falls `case 'e'` through to `case 'E'` (the environment
  display) when `u03`, its unix2003 compatibility flag, is off; the guard reads a dashed `-e` as every
  process, as procps and macOS's default do, since refusing it would refuse every `ps -ef`. The row is the
  recorded gap of the ps item of docs/secret-storage.md.
- The work-budget test gets a second `words` case: 40 keyring execs in front of a 60,000-word tail emit
  2.4 million words at the hops. A mutation control of this round (the word limit ten times larger)
  survived the first version, since the 6,000-hop chain still raised 'words' at 10 million within the 3 s
  bound; the new case does not raise at 10 million (the character budget trips later, in the keyring
  analysis), so the mutant is killed.

Mutation controls of this round: 24 mutants of the guard (a here-document body becoming data again; the ps
union cut to either reading, `-C` a value option on macOS, a dashless cluster at any position on macOS,
a cluster that takes no value, a dashed `-E` on procps; the storage width, each of the four counters and
each of their limits, the mention scan charge, deduplication, main() not refusing a spent budget, the
budget left running, the mention test counting exec arguments, a budget that returns instead of raising;
the systemd-run hint's value-less form), each run against the whole test module in a scratch export with
an 8 s limit per timing row: 24 killed, 0 survived after this commit (23 and the one above before it).

The docs sentence about the work budget no longer names a scratch measurement tool (one word).

Checked: the test module 41 tests, the one failure the tolerated host-copy comparison.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…unt named by its guard

Docs only; the guard is unchanged (sha256 678cd174f3c8bfbe936a8485e21902ed82c2580fd5cc21c4a6b4bbf0a21a9b29,
93,563 bytes), so the SHA256SUMS line stays.

- The loosenings item carries the final differential measured with the guard of 36c847db (the tests changed
  after it, not the guard): 66,176 commands derived from 1,079 table and oracle rows loosen 510, which are
  10 `ps` forms (`ps -fu steve`, `-fu eve`, `-fo user`, `-ft e`, `-fO user`, `-fU steve`, `-fG eve`,
  `-fg steve`, `-fC e`, `-fC eww`) in 51 wrappers each (8 launcher prefixes and 6 suffixes, `bash -c`, `sh -c`,
  `eval`); three seeds of the mutation fuzz (26,400 mutants of the 440 strings the base guard blocks, per
  seed) loosen 153, 175 and 173, and the guard's own segment reading puts every one of those 501 (and the 510)
  in the clustered-value `ps` family, 1,011 of 1,011; nothing is loosened by the 785 fenced blocks, 4,338
  fenced lines, 201 scripts twice, 16,119 distinct repository lines, two fuzzers of 60,000 random strings,
  ten grammar fuzzes of 60,000 launcher chains with redirections, the 40-consumer matrix of 115,200
  commands and the 684 launcher redirection positions. A differential over 478,915 `ps` command lines
  (every one of up to three words from a vocabulary of 65, and 200,000 random ones) loosens 6,673, none
  outside that family.
- The timeout item: 1,000 random mixes of the 22 adversarial building blocks (four seeds) took at most
  0.91 s, and the figures are runs on a shared host (the same 192,000-character quoted word took 0.8 s alone
  and 1.7 s at a load average of 8), so a figure reads within a factor of two.
- The here-document item names the guard its commit-message count was measured with (752def7f, 2,015
  messages then); the final run of this round, 2,065 messages, gives 41 refused by this guard (this round's
  own messages, which name printenv and env in prose, are among the new ones) and 16 by the base guard,
  and none that the base guard refuses passes.

Checked: the four suites 157 tests, the one failure the tolerated host-copy comparison; the whole
repository discovery 7,161 tests, 2 failures (that one and the publication validator, which reports the
pin mismatches of the four changed files, the two files that changed in 172596e without a re-pin, and
nothing else).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The third verification review found that the clustered value-taking ps loosening
lets procps personalities through: with PS_PERSONALITY=old or I_WANT_A_BROKEN_PS
(on the command line or inherited from the shell, which a hook cannot see) procps
parses `ps -axu e` BSD-style, where `u` takes no value and `e` shows the
environment. The coordinator's decision: remove the loosening entirely.

ps_shows_environment now also applies the c26800f rule verbatim
(prior_ps_shows_environment, PS_BSD_CLUSTER): a dashless word of that guard's
cluster alphabet with an `e` is refused wherever it stands, unless a stand-alone
value option takes it. The procps and macOS readings stay for everything else.

Tests: the review's three strings and the brief's rows move to BLOCKED
(environment_dump); the ALLOWED rows that recorded the loosening are gone; the
ps state-machine test adds an independent reference for the c26800f reading.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…er reading refuses

The third verification review and the coordinator's probe found commands that
the base guard (c26800f) refused and this version let through: systemd-run's
--description took the keyring script as its value, so the keyring exec it names
was never unwrapped. A probe of the same class found more: the option values of
run0, systemd-cat and systemd-inhibit, a path-qualified wrapper whose option
value is the script, a systemd-run inside a keyring exec, and a redirection
operator that the base walk took for the value of `sudo -u`, `nice -n` or
`env -u`.

Instead of finding each walk that differs, check() now reads the words twice:
this version's reading first, and when it allows the command, the reading of
c26800f (prior_reading: that guard's walk and rules, ported function by
function, spending from the work budget, with `find -exec` walked by index and
identical started commands read once). A command either reading refuses is
refused, so no command the base refused passes. lex() keeps each text's words
for the call, so the second reading does not tokenize the command again.

Measured in-process against the base guard file (sha256 f9be81b2): over 443,463
commands (test tables in 17 launcher prefixes and 6 suffixes, fenced lines,
19,296 launcher option-value rows, 10 seeded mutants of each of 33,502
base-refused strings) the prior reading equals the base on every one and
check() loosens none. Long env or rtk chains that end in a harmless command now
spend the words budget in the prior reading and are refused as
command_too_complex in under 0.2 s (the base took 29 s and 56 s).

Tests: the four probe strings and eleven more of the class as BLOCKED rows
with the base's reasons; a test that each passes this version's reading and is
refused by the prior one; PATHOLOGICAL rows for the harmless-ending chains.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…; plain text skips shlex

The third verification review found raw STORE_PATHS regular expressions that
run before any work accounting and read an unbounded run again from every
repeat of the literal before it: 102,016 characters of XDG_CONFIG_HOME:- took
11.9 s, HF_HOME:- 13 to 14 s and /proc 13 to 15 s, past the 10 s hook timeout
that fails open. A probe of every text rule found a fourth, the `gh auth
status ... -t` pattern.

Each of the four is now a LinearScan whose search() answers as the old
regular expression: it splits the text once into the runs the pattern cannot
cross (a ${NAME:-default} run of neither `}` nor blank, a path run without a
blank or `//`) and searches each run for literals; the gh rule compares sorted
positions. Old against new: 7,229,043 texts a pattern (2,000,000 random and
every text of up to six fragments), 0 differences; the unit test keeps 100,000
random texts a pattern and every text of up to four fragments.

With the regular expressions linear, one unquoted word of 199,000 characters
still cost 0.53 s in shlex, which copies the word at each character. A text
with no quote and no backslash is now split by SHLEX_PLAIN, which shlex's state
machine reduces to there (its four blanks end a word, a run of its punctuation
is a word, a `#` starts a comment in the legacy reading): 2,111,111 texts in
both readings, 0 differences against shlex.

Measured in-process: the review's three inputs take 0.07 to 0.18 s, and every
STORE_PATHS leading literal repeated to 199,000 characters at most 0.24 s,
before a dump or a harmless command. Tests: those timings (0.5 s bound, the
real hook under 1 s on the review's inputs), both equivalences, and the
backtracking test now repeats each pattern's leading literal as well.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The third verification review measured ('find . ' + '-ok ' * 49000 + '; ...')
at 3.9 s: reader_arguments sliced the rest of the words for every action to
strip its prefix. A prefix that runs to the end from every action
('-exec sudo' repeated, 18,000 times) took more than 20 s however it was
sliced, since every action walked it again.

prefix_end and the prior reading's prior_prefix_end take an optional `ends`
map: a walk from a position ends where the walk from its next step ends, so
the answer is kept for every position a walk steps on and each position is
walked once. Both find readers (this version's and the prior reading's) use it.

Measured in-process: the review's input 0.23 s, its harmless-ending twin (which
reaches the prior reading) 0.32 s, the -exec sudo chain 0.43 s at 330,013
characters. Tests: the four find shapes join the review's timing rows (0.5 s
here, wider in CI; the real hook within a second).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…operator

The third verification review found `set 1 > out` and `set 3 < input` newly
refused (the base guard passed them): the dump rule of set, export, declare and
typeset dropped every redirection from the arguments, and read a number before
an operator as its descriptor because shlex splits `1>out` and `1 > out` alike.
In bash the number is the descriptor only when it touches the operator; apart,
it is an argument (here, a positional parameter).

lex() now marks each number (or bash's `{name}`) that touches `<` or `>`
before it splits the text and hands it back as a Descriptor, a str that equals
the number, so every other rule reads it as before; the dump rule counts only a
Descriptor as a redirection (command_arguments(..., touching_only=True)). A
quoted or glued number is no descriptor, and a mark inside a quoted word is
removed. A text that holds the mark character itself is not marked, and every
number before an operator counts there, as before.

Verdicts against the base guard: `set 1 > out` and `set 3 < input` pass (base:
pass); `set 1>out`, `set >out` and `set 2>/dev/null` print every variable and
stay refused as environment_dump (base: pass).

Tests: the two review rows and four more in ALLOWED, five touching forms in
BLOCKED with the base verdict noted, and a lex() test of the marks against
shlex's words.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e-pass claim is corrected

Under a load average near 20 on this shared host, the harmless-ending find
shape of the last commit took 0.54 s of wall clock against the 0.5 s bound. A
profile showed three Python-level loops that do not depend on the input's
risk: the default-run scans (XDG_CONFIG_HOME, HF_HOME) and the /proc scan
walked every run of the text (49,000 there), each find action walked its
prefix even when the next word is an option, and is_environment_dump built
the argument list of every segment. Now:

- default_run_scan and proc_environ_scan read only the runs that hold their
  literal, from each first occurrence to the end of its run (the same answers:
  7,229,043 texts a pattern against the c26800f regular expressions, 0
  differences);
- a find action followed by an option word (dash-led, no slash) is skipped in
  both readings: the walk would stop on that word, which is no reader;
- is_environment_dump reads the arguments of set, export, declare and typeset
  only.

Measured in-process: the harmless-ending find shape 0.11 s (was 0.32 s), every
STORE_PATHS leading literal at 199,000 characters at most 0.12 s. The timing
test bounds processor time at 0.5 s (wall clock 3 s), since the host is shared.

The module docstring no longer says that every scan reads a text once (the
third review showed four STORE_PATHS patterns that did not); it names the
LinearScan rules, the plain split, the descriptor mark and the two readings.
The WORK_LIMITS comments say both readings spend from one budget and give the
2026-09-30 measurements. The oracle table comment counts the added rows as the
review asked: 245 of the 342 added BLOCKED and KEYRING_BLOCKED rows were
allowed by the base guard, 97 are regression controls.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…iptor adjacency, recount

docs/secret-storage.md, guard section, dated 2026-09-30 statements:

- The opening paragraph and the former "Loosenings" item: the clustered-value
  ps loosening and the claim that it was the only block-to-allow difference
  are withdrawn; the new "No loosening" item names the third review's
  launcher strings and the class the probe found, the prior reading that
  closes it by construction, the evidence (the prior reading equals the base
  guard on 460,910 commands with the final guard; no harness loosens any
  command), the 44 reason changes that 6c4f63d already had, and the recount
  of added rows (245 new refusals, 97 regression controls of 342).
- The ps item: personalities, the base guard's reading kept, the forms that
  are refused again, and agent-safe forms checked against this guard
  (`ps -f -U 1000`, `ps -fU 1000`, `pgrep -a -u steve node`).
- The redirection item: a number is a descriptor for the dump rule only when
  it touches its operator (`set 1 > out` passes, `set 1>out` stays refused).
- The timeout item: the one-pass claim is corrected (four STORE_PATHS patterns
  were quadratic), with the linear scans, the plain split, the find walk, the
  prior reading's chain limits and the re-measured budget and friction figures
  (sh -c strings 4,999; two commit messages of the refs over the characters
  budget, as with 6c4f63d).

The guard's WORK_LIMITS comment gives the mix timing as measured over several
runs (a comment-only change: the AST equals the guard the harnesses ran on);
SHA256SUMS carries its hash.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ot-file commit, last)

Hashes of the six changed pinned files on top of main's manifest at b4056a3 (docs/lanes.md protocol: main's copy, register each
file, component_matrix and new_host_grand_list --write), plus the receipt guard-k3-verification-20260930 (artifact
measurement: the acceptance oracle, probes, the documented-command replay, real hook timings and the two GPT-6
verification rounds of this change); validate.py passes: 69 components, 8297 hashed files, 175 receipts. The installed
host copy is reinstalled after merge and after Gate A window W, announced first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

lane:foundation Foundation lane: Claude/Codex setup, hosts, memory, RAG, research, workers

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant