Repository navigation
Command guard tightening: systemd-run launchers, substitutions in double quotes, macOS ps -E, systemctl show-environment, fail-closed and work budget - #511
Merged
Conversation
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
seathatflowsinourveins
force-pushed
the
claude/guard-launchers-20260929
branch
4 times, most recently
from
October 1, 2026 03:19
dc33b48 to
fb4a586
Compare
4 tasks done
… command line Failed first: 30 of 31 new BLOCKED rows in tests/test_secret_path_guard.py returned None instead of their reason (systemd-run --user --pipe --wait cat "$PAPER_ENV_FILE" and its bash -ic form, printenv and env behind it, an option's value taken for the command, the -E/--setenv/-p Environment= secret names), with 60 more failures behind rtk proxy and 23 behind a keyring exec; the oracle showed 8 systemd-run MUST_BLOCK mismatches (22 in all). The 7 new ALLOWED controls (the trading lane's loader path and ordinary units) passed before and after. Design: expand() unwraps systemd-run like env and rtk. It skips the launcher's own options (wrapper_options, from the getopt table of systemd v255 src/run/run.c, plus the value options of v256-v258 so a newer host still finds its command), keeps the systemd-run segment in the result, and reads the started command with every rule, a nested bash -ic string too. segment_reason() blocks a secret variable NAME (SECRET_NAMES) set through -E/--setenv or -p/--property Environment=, with or without a value, as the new reason secret_variable_on_command_line: the command line lands in the journal (_CMDLINE) and the unit's properties travel over the user bus. skip_wrapper_options now delegates to wrapper_options. The two EXPECTED_PASS_THROUGH rows that recorded the gap moved to BLOCKED. Checked: oracle 22 -> 14 mismatches (none left for systemd-run); tests.test_secret_path_guard green except the host-copy comparison, which is fixed by reinstalling the guard after merge. Differential against b40b359 over 48,310 generated commands (every table row wrapped in launchers and suffixes): 0 loosened, 0 reason changes without the new launcher, 0 newly blocked commands without it. Randomized comparison of the refactored option walker with the base one: 400,000 cases, 0 mismatches. SHA256SUMS carries the new guard hash; the guard section of docs/secret-storage.md records the rule. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Failed first: 24 of the 26 new BLOCKED rows in tests/test_secret_path_guard.py returned None
instead of their reason (echo "$(printenv)", echo "`printenv`", x="$(printenv)"; echo "$x",
git commit -m "$(printenv)", a backticked `set` in prose inside double quotes, nesting in
either order of quoting, a substitution inside ${x:-...} and $(( ... )), and the credential
file, .env, secret-name, keyring, native-token and trace rules reached through a body), with
48 more failures behind rtk proxy and 23 behind a keyring exec; the oracle showed 6
MUST_BLOCK mismatches for these forms. The 2 rows that already passed are echo "a $(echo
"$(printenv)") b" (the base tokenizer's quote parity happened to expose the inner body) and
a curl row whose reader already reads ~/.ssh. A second red step: 5 of 9 rows of the scanner
table for here-documents failed before the scanner learned to skip their bodies (below).
Design: substitution_bodies() scans the command once with a small frame stack (unquoted,
double-quoted, `$(` body, backquote body) as the Bash Reference Manual describes it: double
quotes keep $ and the backquote special, a backslash escapes only $ ` " and \, single quotes
are data outside double quotes, a `$(` body ends at its matching parenthesis and a backquote
body at the first unescaped backquote. It returns the outermost bodies that start inside
double quotes; expand() reads each one as a command, level by level to 32 levels, so every
rule and every wrapper (rtk, keyring exec, systemd-run, nested shells) applies to it. An
unquoted $( ... ) stays with the tokenizer, and single-quoted text and a backslash-escaped $(
or backquote stay data (the oracle's MUST_ALLOW rows and 12 new ALLOWED rows pin them). The
scanner treats a here-document's body as literal text, as bash's parser does when it looks for
the `)` that ends a `$(`: prose in a commit message (an unbalanced parenthesis, an apostrophe,
backquotes) opens nothing. Without that, one commit message produced 29 spurious bodies. How
the guard reads here-document bodies as commands is untouched (out of scope).
Checked: oracle 14 -> 8 mismatches (the rest belong to ps -E and systemctl). Table test of 30
texts against their expected bodies, and a bounded-nesting test (5000 levels neither raise nor
hang). tests.test_secret_path_guard green except the host-copy comparison. Differential
against b40b359 over 50,178 generated commands: 0 loosened, 0 reason changes without a new
form. Plain-form equivalence, the promise of this change: for 516 single-line table and oracle
rows in three shapes ("$(ROW)", "`ROW`", X="$(ROW)"), the new guard's verdict on the
double-quoted form is never more lenient than the base guard's verdict on the unquoted
substitution (0 of 1,548) and never differs in reason; it is stricter for the systemd-run rows
(that launcher was unread in both forms) and for one row where shlex fuses `$(>` into one
token in the plain form. Real corpus: 16,089 distinct lines and blocks of this repository's
Markdown fences and shell scripts, base blocks 33, this change 34, 0 loosened; the one new
block is a true positive (env | cut -d= -f1 inside a double-quoted substitution). Real commit
messages, 840 of them in the git commit -m "$(cat <<'EOF' ... EOF)" pattern: the new verdict
equals the base verdict on the unquoted form for 839 (6 -> 8 blocked; the 2 new blocks are a
prose line starting with `set` that the base already flags in the unquoted form, and item 1's
own message quoting cat "$PAPER_ENV_FILE" behind systemd-run).
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…onment dumps
Failed first: 14 rows in tests/test_secret_path_guard.py returned None instead of
environment_dump (ps -E, -Ewwp 123, -p 123 -E, -A -E, -AE, -eE, -ef -E, -o pid,command -E,
Eww 123, auxE, E, and three keyring-exec forms), with 22 more failures behind rtk proxy and
11 behind a keyring exec; the oracle showed the 6 ps MUST_BLOCK mismatches. The 7 new ALLOWED
controls (ps aux, -ef --sort, -eo pid,etime,args, -o pid,ETIME, -u Eve, -C E, -o etime=)
passed before and after.
Why: macOS documents -E as the environment display, "-E Display the environment as well", and
lists the BSD-style e as "Same as -E" (Apple adv_cmds ps.1, read 2026-09-29). The guard read
only dashless clusters with a lower-case e. A dashed -e is every process on Linux and macOS
("Identical to -A"), so ps -ef stays allowed.
Design: PS_BSD_CLUSTER accepts E as well as e, in a strict superset of the old language, and
ps_shows_environment() reads a dashed word's letters up to the first option that takes a value
(PS_ARG_OPTIONS), so -Ewwp 123 and -p 123 -E are found while -pE, -uE and -u Eve, where the E is
a value, are not. The dashless branch and the value-skipping are unchanged.
Checked: oracle 8 -> 2 mismatches (the two systemctl rows); tests.test_secret_path_guard green
except the host-copy comparison. 20,000 random realistic ps lines against b40b359: 0 loosened,
5,181 newly blocked (3,899 distinct) and every one carries an E flag; capital-E values and
names still pass.
Differential over 51,174 generated commands: 0 loosened, 0 unexplained reason changes, 0 newly
blocked commands without a new form. Real corpus of 16,089 repository lines: no new block.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…nment dump Failed first: all 18 new BLOCKED rows in tests/test_secret_path_guard.py returned None instead of service_manager_environment (systemctl [--user] show-environment, options before and after the verb, -M/-H and --machine/--host values, `--`, a redirection between options and verb, sudo, a path to systemctl, timeout, a pipe, bash -c, systemd-run --pipe, and inside a double-quoted substitution), with 36 more failures behind rtk proxy and 18 behind a keyring exec; the oracle showed the 2 systemctl MUST_BLOCK mismatches. The 9 new ALLOWED controls (show -p Environment UNIT, cat, status, list-units, is-active, show -p MainPID --value UNIT, -H host status, restart, and a unit named show-environment.service) passed before and after. Why: systemctl(1) 255 says show-environment dumps "the systemd manager environment block. This is the environment block that is passed to all processes the manager spawns", so any credential that was imported into that manager lands in the output, as with env. Design: systemctl_verb() reads systemctl's first word after its options, using the option table of src/systemctl/systemctl.c at systemd v255 (getopt string "ht:p:P:alqfs:H:M:n:o:iTr.::", no leading `+`, so options may also follow the verb; the value options are -t -p -P -s -H -M -n -o and the 25 long ones of v255, plus the v256-v258 additions) through the same option walker as systemd-run, with redirections dropped first. segment_reason() returns the new reason service_manager_environment for the verb show-environment. It is not `environment_dump`, so the keyring-exec test requires the exact reason, and the rule still applies behind a keyring exec and rtk proxy. HINTS names the allowed alternative, systemctl show -p Environment UNIT. Checked: oracle 2 -> 0 mismatches, exit 0 (65 cases); tests.test_secret_path_guard, tests.test_credential_status, tests.test_credential_tools and tests.test_install_claude_profile green except the host-copy comparison, which is fixed by reinstalling the guard after merge. Differential against b40b359 over 52,672 generated commands: 0 loosened, 0 unexplained reason changes, 0 newly blocked commands without a new form. Real corpus of 16,089 repository lines: no new block from this change. Known gap, recorded in docs/secret-storage.md: systemctl show with no unit prints the manager's own properties, Environment= among them. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… command, an internal error blocks Failed first (tests/test_secret_path_guard.py, run against the previous tip 03d3871): 7 of 9 PATHOLOGICAL timing rows failed, each run in a child process with a 45 s limit and a 3 s bound on check() itself: 12,000 here-documents 10.6 s, 12,000 distinct delimiters 10.3 s, 12,000 lines of $((1 << 2)) 10.1 s, 20,000 nested env 28.3 s, 20,000 nested rtk proxy and 60,000 nested systemd-run past 45 s, and the 20,000-deep quote nesting returned environment_dump where the row expected the base verdict (its expected verdict is environment_dump: shlex's quote parity exposes the innermost command one level down). 10 more rows failed: 8 substitution_bodies rows for $((...)) and the allowed row env=2; echo "$((env))", and 3 error-handling rows (main() raised instead of blocking). A hook that times out does not block the call (Claude Code hooks documentation, "Timeouts", read 2026-09-29), so a slow guard fails open. Design: substitution_bodies() is one pass over the text with a stack of frames (unquoted, double-quoted, $( ), backquotes, ${ }, $(( )) ) and a regular expression that jumps to the next character that matters, instead of stepping through every character. A here-document's terminator is found with a binary search in a line index built once per text (a rescan per << was the quadratic part), and an unterminated one is no here-document, as before. $(( is arithmetic only when a "))" that touches closes it, so $((env)) reads a variable, a << in it is a shift, real substitutions inside it are still read, and $((printenv) ) is a substitution holding a subshell. launcher_chain() walks env, rtk and systemd-run hops with an index instead of copying the rest of the words at each hop (prefix_end() replaces the copying strip_prefix loop); only a systemd-run hop keeps a segment of its own, because no rule reads an env or rtk hop that starts a command. expand() stops reading nested bodies at 32 levels or at four times the command's length plus 64 KiB. main() catches any exception from check() and blocks with one line, "blocked (guard_error)", without command text or traceback (new HINTS entry); unreadable hook input still exits 1 and other tools exit 0. Checked: the nine PATHOLOGICAL rows now take 0.11 to 0.75 s (test bound 3 s); before/after on this host, seconds: 12,000 here-documents 10.4 -> 0.16 (base 0.14), 12,000 shifts 10.1 -> 0.24 (base 0.22), 6,000 "cat <<X" 2.6 -> 0.11, 20,000 nested env 28.1 -> 0.11 (base 29.1), 20,000 nested rtk 55.1 -> 0.16 (base 55.9), 60,000 nested systemd-run > 60 -> 0.75 (base 0.55, which did not read it), 1,600 nested systemd-run 9.98 MiB -> 0.23 MiB (base 0.17 MiB), 20,000-deep quote nesting 4.9 -> 0.72. A realistic 1 MB command is bounded by shlex, unchanged and not part of this work: a 840 KB git commit -m "$(cat <<'EOF' ...)" takes 8.5 s at base, 8.6 s at 03d3871, 9.0 s now. Oracle: 65 cases, 0 mismatches. Suites: only test_host_profile_copy_is_verbatim fails. Differential against base c26800f over 52,960 generated commands: 0 loosened, 0 unexplained reason changes, 0 newly blocked without a new form; against 03d3871 the only differences are 65 $((...)) forms that pass again (base passed them). Walker comparison with base, 400,000 random cases: 0 mismatches. Random grammar fuzz, 120,000 strings against base: 0 loosened, 0 exceptions. Real repository corpus (16,119 lines and blocks): base blocks 33, now 34, the same single true positive as before, 0 loosened. tests: check() runs without raising over 834 rows: BLOCKED 303, KEYRING_BLOCKED 127, ALLOWED 119, SAFE_CORPUS 139, EXPECTED_PASS_THROUGH 36, the four oracle groups (now copied into the test file) 65, the 36 scanner-table texts and the 9 pathological inputs. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ere-documents are data, ps -fu Eve passes Failed first (tests/test_secret_path_guard.py, run against the guard of 2f097bf): 64 subtest failures on the new rows, none of them passing by luck. Blocked before and now allowed or blocked as they should be: 11 BLOCKED rows returned None ("# macOS" then ps -E, "# check ..." then systemctl --user show-environment, "# run it" then systemd-run --pipe --wait printenv, echo ${#PATH}; printenv, echo $#; printenv, gh api .../issues/1#c; cat "$PAPER_ENV_FILE", echo a#b; printenv, two commands with a trailing comment, echo "$(date # )" whose printenv line follows the comment, a comment holding quotes and a backquote, and a here-document heading followed by printenv), plus 22 behind rtk proxy and 11 behind a keyring exec. 8 ALLOWED rows returned a reason: echo ok # "$(printenv)", printf '%s' $'it\'s "$("printenv")"', ps -fu Eve (two rows) and four commit-message patterns with a quoted here-document. 11 scanner-table rows and 1 expected-pass-through row (a shell reading a quoted here-document in a substitution) differed, and all 10 subtests of the five real commit messages (cf58426, 32a6cd2, a84fa7c, 88af2ba, b553826, verbatim, in the standard git commit -m "$(cat <<'EOF' ...)" and the gh pr create --body forms) failed. The bug behind the first group is older than this work: tokenize() joins lines with ";" and shlex reads a "#" anywhere, "$#" and "a#b" included, as a comment, so a "#" dropped the whole rest of the command (found by the 2026-09-29 review). Design: scan_shell() (substitution_bodies() now calls it) returns the bodies and the comment spans from the same single pass. A comment is a "#" that starts a word (at the start of the text or after a blank or one of ;&|()<>, and not right after a quote, escape or substitution), outside quotes and outside a double-quoted substitution, to the end of the line; inside a $( ) body it also runs to the end of its line, so a ")" in it closes nothing, and in a backquote body it ends at the closing backquote. tokenize() removes the spans and sets shlex's commenters to "", and command_segments() also reads the command as shlex did before (for commands up to 200,000 characters) and adds only the segments that reading lacks, so nothing the guard read before is dropped. In a here-document body a comment hides the rest of that body, as before, not the text after the terminator, so a script written through a here-document (its first line is #!) stays as unread as it was. $'...' is data up to the first quote that no backslash escapes, except inside double quotes. A here-document whose delimiter is quoted (also partly: <<'E'OF) is data: the body that scan_shell returns loses it, keeping the operator line and the terminator; an unquoted delimiter keeps its body, as at the top level. ps_shows_environment() skips the word after a cluster that ends in a value option (ps -fu Eve). The three vacuous BLOCKED rows are marked as regression rows and two are replaced by rows that pass at base, echo "$(echo \"$(printenv)\")" and echo "$(base64 ~/.ssh/id_ed25519)". Checked: oracle 65 cases, 0 mismatches; suites: only test_host_profile_copy_is_verbatim fails (139 tests). Differential against base c26800f over 54,126 generated commands: 0 loosened, 0 unexplained reason changes, 0 newly blocked without a new form. Random grammar fuzz against base, 160,000 strings in 4 seeds: 0 loosened, 0 exceptions. Real repository, by realistic shape: 785 fenced blocks, 4,338 fenced lines and 201 scripts written through a quoted here-document: 0 newly blocked; the synthetic shape of a whole script passed as one command string newly blocks 4 of 201, none of them a true positive and all long-standing readings that the old "#" had hidden (three array literals x=(env -i ...) and a python set(...)). Commit messages: the standard pattern newly refuses 0 of 842 (both guards refuse the same 8; the previous tip newly refused 3 of the 842 and 6 of 1,865 across all refs); the top-level git commit -F - shape newly blocks the two messages of this work that quote the new ps -E and systemd-run rules in prose, by those rules and not by the comment change. 1 MB timing: a 840 KB git commit -m "$(cat <<'EOF' ...)" takes 8.5 s at base and 8.9 s now (two readings would take 17.8 s, so commands above 200,000 characters are read once). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ocuments, systemctl show, quoted wrappers), correct the claims Failed first (tests/test_secret_path_guard.py, run against the guard of 0ad4835): 154 subtest failures on the new rows, each a reason expected and None returned. 41 BLOCKED and KEYRING rows, plus 78 behind rtk proxy and 31 behind a keyring exec (numbers before the leading-redirection rows, which failed the same way): a backquote in a single-quoted string handed to bash -c, sh -c and eval; redirections between systemd launcher options (systemd-run --user 2>/tmp/log --pipe printenv, ... > /tmp/log -E HF_TOKEN true); $(<.env), $(< ~/.aws/credentials), x=$(<~/.netrc), echo `<.env`; unquoted here-document bodies that run $(...) (cat <<EOF with value: "$(printenv)"); run0, systemd-inhibit and systemd-cat; a bare systemctl show and show -p Environment; /usr/bin/sudo, /usr/bin/timeout, /usr/bin/nice; ps -CE and ps -C -E; systemctl and systemd-run behind watch and flock inside a keyring exec (the wrapper matrix rows). 3 scanner-table rows for here-document bodies failed too. The 19 new ALLOWED controls (the alternatives the docs and the hint recommend, $(<version.txt), a quoted here-document, path-qualified launchers of ordinary commands, ps -C python3, systemctl show naming a unit) and the 12 new expected-pass-through rows passed before and after, as they must. Design: scan_shell() also returns the backquotes that sit in single-quoted and ANSI-C strings (tokenize() keeps them, so the shell that string is handed to sees them: item 9), the bodies of an unquoted here-document scanned as double-quoted text with no closing quote (a "hd" frame, one substring scan per body, nested at most 8 deep), and the quoted here-document bodies inside a substitution that closes, which the tokenizer no longer reads in either of its readings (the five real commit messages, and two commit messages of this work that the earlier guard refused through shlex's quote parity, now pass in the standard pattern). wrapper_options() skips redirections for the systemd launchers only: for timeout the duration in "timeout 5 > out cmd" reads as a descriptor and the command was swallowed (the base differential caught it), so every other wrapper keeps the walk it had. A token "(<" (shlex joins punctuation that touches) and a segment that is only "< FILE" add a "cat FILE" segment, and a leading "< FILE cmd" adds "cmd < FILE", in addition to what was read before. The systemd launcher tables (run0 from parse_argv_sudo_mode v256-v262, systemd-inhibit, systemd-cat) and the value options of later systemd-run releases cite the tag that first has each (--root-directory v259, --output v261: the review found them absent from v256-v258). systemctl_call() reads the verb and its arguments wherever the options stand, so a bare show (or show -p Environment) with no unit is refused. prefix_end() compares the program name of a wrapper, so path-qualified launchers count. ps_shows_environment() refuses an E word right after -C. launched_commands() reads systemctl and the systemd launchers, and keyring_reason() returns service_manager_environment for a started systemctl. The earlier reading of a command that holds a # (or a protected backquote, or data) is kept beside the new one, for commands up to 200,000 characters, so nothing the guard read before is dropped. The dated docs subsection is rewritten with the corrected claims: the manager's "Started <unit> - <command line>" journal line (measured on systemd 255.4, not _CMDLINE); --pipe returns the output, --wait shows terse unit information and the output goes to the journal without either; the scope of "every rule applies"; the alternatives considered for reading shell syntax (bashlex, tree-sitter-bash, mvdan/sh, with the dates and licences read from PyPI and the GitHub API on 2026-09-29); the residuals with their inert strings; and the systemctl hint now names the alternatives. Three files that recommended the refused command (omniroute.service, token-report-refresh.service, tools/token-report/README.md) keep it and add that it is for the operator's own terminal, with the agent alternative. Checked: oracle 65 cases, 0 mismatches; suites (4 named): only test_host_profile_copy_is_verbatim fails (141 tests); the suites that read the three edited files (omniroute unit, token-report units, ecosystem manifest, dashboard data, workflow hardening, adoption docs and status, 319 tests) pass. Differential against base c26800f over 59,530 generated commands: 0 loosened, 0 unexplained reason changes, 0 newly blocked without a new form. Random grammar fuzz against base, 240,000 strings in 6 seeds: 0 loosened, 0 exceptions. Mutation fuzz of strings the base blocks (quotes, comments, backquotes, here-documents and launchers inserted), 8 seeds of 22,440 mutants: 1 differs (the run before the last change, seeds 21 to 28, had 3), every one a quoted here-document whose body holds the printenv (bash runs nothing there). Walker comparison with base, 400,000 random cases: 0 mismatches. Real repository by shape: 785 fenced blocks, 4,338 fenced lines and 201 scripts written through a quoted here-document newly block 0; a whole script passed as one command string newly blocks 4 of 201 (array literals x=(env -i ...) and a python set(...), read as commands as before). Commit messages: the standard pattern newly refuses 0 of 1,884 (all refs); two messages of this work that the earlier guard refused there now pass; the unquoted-delimiter and the top-level git commit -F - shapes read prose as commands as before (newly refused 8 and 4 of 843, by the # fix and by this work's own new rules quoting themselves in prose). tests: check() runs without raising over 968 rows. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ommand takes data The data rule of 172596e treated the body of a quoted here-document inside a double-quoted command substitution as data whatever received the substitution. The coordinator's probes found the cost: bash re-reads what a substitution prints as code for `eval "$(...)"`, `bash -c "$(...)"` (also sh, zsh, behind env, timeout, nohup, xargs) and a code consumer nested in an echo's substitution, and the base guard refused all of those as environment_dump while 172596e allowed them. Failed first (the new tests against 172596e): 196 failures and 2 errors, one failure being the tolerated host-copy comparison. The 38 new BLOCKED rows failed directly (38), behind rtk proxy (76) and behind a keyring exec (38); 27 consumer rows, 13 substitution-body rows and 2 oracle groups failed; the two data-rule tests errored on the missing names. Design: scan_shell() decides each quoted here-document only after the whole text is read. data_consumers() reads the text once more with every returned substitution replaced by a marker and the comments and top-level here-document bodies cut, finds the command each marker is an argument of after the reserved words, assignments, wrappers and launchers the guard models (receiving_command), and answers True only for git, gh, echo, printf, cat and tee. A shell, eval, an interpreter, source, xargs, watch, ssh or any other program, an assignment, the command position, a marker before the program word, a substitution nested in another, and any text the guard cannot tell (markers in the text, more than 100,000 characters left after the cuts, more than 16 launcher hops) keep the old reading: the body read as command lines and the tokenizer input whole. Only a here-document at the command level of the returned body is data, so `echo "$(echo "$(cat <<'EOF' ...)")"` keeps the strict reading. cat and tee count only for a here-string word (`cat <<< "$(...)"`): a matrix of 39 heads x 18 launcher prefixes x 5 here-document forms x 8 payloads x 4 substitution forms (112,320 commands) showed the plain allowlist loosening `cat "$(cat <<'EOF' ... a credential path ...)"`, which the base refused through the reader rules, because a reader's operand is a file name; the operand forms now keep the base verdicts. Checked: kw/guard_probe_final.py base candidate reports loosened rows: 0 (172596e: 1); the eight-consumer probe, with rtk proxy and keyring exec forms, 0 of 13 loosened (172596e: 13); oracle 65 cases, 0 mismatches; the matrix loosens 0 rows for every head that is no data consumer (172596e: up to 2,040 per head) and 0 for git, gh, echo and printf; mutation fuzz, 60 mutants per each of 397 blocked strings: 6 loosened at seed 7 and 2 at seed 11 (172596e: 387 and 378), each a mutant whose broken terminator or quote makes the rest of a git or echo substitution here-document data that bash never runs, or a bash syntax error; differential against c26800f over 59,592 commands from 936 rows: 0 loosened; the real-repository corpora (fence lines and blocks, scripts through here-documents): 0 loosened; tests.test_secret_path_guard 29 tests, the host-copy comparison the only failure. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…and_too_large)
Claude Code's hooks documentation ("Timeouts", read 2026-09-29) says a timed-out command hook
does not block the tool call, and this hook's timeout is 10 s. Measured on this host on
2026-09-29 with one quoted word: 1,000,000 characters took 10.2 to 13.0 s in check() (base
c26800f and this version alike, over several runs), 600,000 took 3.8 to 4.4 s and 500,000 took
2.9 to 3.3 s, so a large enough command outlasts the hook and runs unread.
Failed first: test_the_size_limit_is_at_600000_characters and
test_a_command_over_the_size_limit_is_refused_without_echoing_it errored on the missing
MAX_COMMAND_CHARACTERS and the missing reason. The third test, a 500,000-character ordinary
command that must not be refused for its size, passed before and after, as it should.
Design: main() refuses len(command) > MAX_COMMAND_CHARACTERS (600,000 characters, not bytes)
before any rule reads the command, with exit 2 and one stderr line that names no command text:
"blocked (command_too_large)" and the hint to put the content in a file with the Write tool and
pass the path. check() is unchanged for size, so every table row behaves as before, and a
command of exactly 600,000 characters still reaches check().
Checked: a 700,000-character command is refused in under 5 s through the real hook subprocess
(exit 2, one line, neither the marker text nor its filler on stderr or stdout); 500,000
characters of ordinary shell (quoted words, arithmetic, redirections, separators) exit 0 with no
output inside 5 s and, with printenv appended, exit 2 with environment_dump; the boundary
test passes 600,000 characters to check() (mocked) and refuses 600,001 without calling it; 300,000
four-byte characters, 1.2 MB of UTF-8, pass. tests.test_secret_path_guard: 32 tests, the
host-copy comparison the only failure.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ries the new guard docs/secret-storage.md, guard section, dated 2026-09-29 and no restructuring: one new item for the rule that a quoted here-document is data only for git, gh, echo, printf, cat (as a here-string word) and tee. It says why (bash runs what a substitution prints for a shell with -c, eval, source, xargs or an interpreter), that every unlisted consumer keeps the stricter reading so prose in a quoted here-document inside them can still be refused, what the guard still does not read (an allowed consumer's output used later, a shell reading the quoted here-document itself), and the measured matrix. The here-document item now says "where the command takes data", and the timeout item gains one sentence for command_too_large with the measured timings (1,000,000 characters of one quoted word: 10.2 to 13.0 s; 600,000: 3.8 to 4.4 s; 500,000: 2.9 to 3.3 s), replacing the "not fixed" note that described the gap. The guard's own docstring says the same in two sentences. adoption/hooks/claude/SHA256SUMS: only the guard's line changed, to 914342a040d9632b9aedc435e58617a78808b5f5f6fcbe88a9460edda41be75c (91,845 bytes); sha256sum -c over the file reports every line OK. Checked: kw/guard_oracle.py exits 0 (65 cases, 0 mismatches); kw/guard_probe_final.py reports loosened rows: 0 with the timing rows of the last round unchanged (12,000 here-documents 0.19 s, 20,000-deep quoting 0.66 s, 60,000 systemd-run 0.58 s); tests.test_secret_path_guard, test_credential_status, test_credential_tools and test_install_claude_profile ran 148 tests, the host-copy comparison the only failure; 305 further tests (docs consistency, adoption status, blind checkout, effort guard, release pins, landscape) pass; scripts/validate.py lists only the manifest hash and byte mismatches of the pinned files this branch changed (SHA256SUMS, the guard, its tests, docs/secret-storage.md, and the two files 172596e changed without a re-pin, omniroute.service and tools/token-report/README.md). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…e inventory entry) A later inventory entry, canary-e2e (a disposable synthetic proof key: class test_canary, status test_only), declares the variable CANARY_E2E_KEY, and test_inventory_secret_names_are_all_guarded requires every inventory name to be in SECRET_NAMES, so it would fail there without the name. It is added now, ahead of that entry (which belongs to another branch and is not touched here), and is refused like every other name. Failed first (the two new rows against 914342a0, the guard of 82767ca): 9 failures, one being the tolerated host-copy comparison; the rows failed directly (2), behind rtk proxy (4) and behind a keyring exec (2). The four forms `echo "$CANARY_E2E_KEY"`, `echo $CANARY_E2E_KEY`, the os.environ lookup and `rg -n CANARY_E2E_KEY` all returned None; they now return secret_variable_reference (three) and secret_name_search, and `rg -n MY_CANARY_E2E_KEY_HINT` still passes (the word boundary). Design: one entry in SECRET_NAMES with a comment; the expansion, lookup and search rules and test_every_secret_name_is_caught_by_a_search pick it up from the tuple. Explicit BLOCKED rows for `echo "$CANARY_E2E_KEY"` and `rg -n CANARY_E2E_KEY`, so the wrapper tests also run them behind rtk proxy and a keyring exec. No docs sentence lists where the names come from (only the comment above the tuple does), so the docs are unchanged. The inventory tie holds with and without the later entry's variable (29 inventory names, 30 unique guard names, none missing). SHA256SUMS carries the final guard hash for this and the earlier changes of this round. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The independent verification review asked for a limit low enough that the tokenizer's two readings of a command (with and without shlex's comments) always both run, so that no reading is skipped above some length. Measured on this host on 2026-09-29, one quoted word costs 0.6 s at 200,000 characters (about 1.5 s for both readings), 2.9 to 3.3 s at 500,000, 3.8 to 4.4 s at 600,000 and 10.2 to 13.0 s at 1,000,000 (base c26800f and this version alike), against the hook's 10 s. Failed first: the size test that pins the limit failed (600000 != 200000) once it asked for 200,000; the others use the constant and follow it. Tests: a 250,000-character command is refused through the real hook (exit 2, one line, no command text) inside 5 s; a 150,000-character ordinary command exits 0 with no output inside 5 s and, with printenv appended, exits 2 with environment_dump; the boundary test passes 200,000 characters to check() (mocked) and refuses 200,001 without calling it; 150,000 four-byte characters (600 KB of UTF-8) pass. The constant's comment, the guard docstring and the docs sentence carry the new limit and the timings; SHA256SUMS carries the guard's new hash (a stale one fails four tests of the other suites). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…acktracking regular expression The independent verification review found that `'ps ' + 'E' * 70000 + 'q; printenv'` took 13.2 s in check() against 0.08 s at the base guard: PS_BSD_CLUSTER, `^[aAcfhjlmrsStTuvwxXLnE]*[eE][aAcefhjlmrsStTuvwxXLnE]*$`, has two overlapping quantifiers that fail in quadratic time, and a hook that outlasts its 10 s timeout does not block the call. Reproduced here at 15.2 s. Failed first: the new timing row "ps cluster of 70,000 E and a letter that is no flag" failed at 15.2 s against the 3 s bound, and the equivalence test errored on the missing function. Design: is_ps_bsd_cluster(word) is PS_BSD_LETTERS.issuperset(word) and an `e` or `E` in it, the same language as the regular expression (checked on every word of up to 5 letters over nine letters, 66,430 words, and on 13 words the guard meets). Five timing rows for ps clusters of 70,000 characters (a lead of E, of e, of a, a trailing E, a letter that is no flag) join the PATHOLOGICAL table with their verdicts and the 3 s bound, and a test times every compiled pattern of the guard (module level, the scan and store tables: 52) on 70,000 repeats of 12 characters with a lead and a tail that make a match fail late, each under a quarter of a second. Negative control: the same test against the guard before this change fails naming PS_BSD_CLUSTER at 14.6 s. The audit script that found nothing else (PAREN_INPUT is slow under search, 6.4 s, but the guard only calls it as fullmatch, which is linear) also timed 2,132 check() shapes, a first word (ps, env, sudo, systemd-run, systemctl, keyctl, rtk, grep, git, bash -c ...) followed by 70,000 repeats of a character or a token: the slowest were the two ps shapes, everything else 0.3 s. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…the only exemption The independent verification review of 172596e (GPT-6, effort max, read-only) found six blocking block-to-allow regressions, five in the general here-document machinery: a quoted delimiter made the body of a here-document inside a double-quoted substitution data, and each imperfection of the bash-exact parsing that needs erased executable text. Reproduced here as inert strings, each an environment_dump at the base guard c26800f and None at 172596e: `echo "$(bash <<'EOF'` with `echo "$(printenv)"` in its body (a shell reads the body); a `<<$'EOF'` delimiter; a body line `text\` that the guard's backslash-newline join glued to its terminator; an arithmetic command `((1 << "2"))` taken for a `<<`; a quoted `#` (`printf ' #x'`) taken for a comment that hid the rest of a here-document body. The design change of the coordinator replaces the general reading. Failed first (the new tests against 3713250): 125 failures and 129 errors, one failure being the tolerated host-copy comparison: the recognizer table (61 rows), the consumer table (60 rows), the 41 new BLOCKED rows directly (29), behind rtk proxy (58) and behind a keyring exec (29), the plain-scan rows of the bodies table (6), the ANSI-C rows (7), a known-bypass row, the neutral reading and the joined-line test. Design. scan_shell() no longer reads here-documents at all: the delimiter regular expressions, heredoc_delimiter, heredoc_end, the terminator index in the scanner, the here-document data spans, the comment-in-a-body rule and the `hd` frame are deleted, so a here-document's lines are command lines as the tokenizer has always read them at the top level (a `#` hides its own line, a `"$(...)"` in a body line is a substitution). The one exemption is idiom_spans(): a double-quoted word that is exactly `"$(` blanks `cat` blanks `<<` [`-`] blanks `'IDENT'` blanks newline, the body, the FIRST line that is exactly IDENT (for `<<-` after its leading tabs, no other trimming, no carriage return), blanks and newlines, `)"`, standing alone as a word. It is looked for in the text as written, before check() joins backslash-newline pairs, with a line index built once and a tail memoised per terminator, so no head costs a rescan. exempt_idioms() then reads the text once more with each idiom replaced by a marker and asks the command that receives it: git, gh, echo or printf after the reserved words, assignments, wrappers and launchers the guard models, with the idiom a word of its own after the program, not inside the body of a substitution, and the text readable by shlex without its fallback. Then neutral_reading() puts a neutral word in its place for the COMMAND reading only; the raw-text rules (store paths, secret names and expansions, /proc) read the whole text as before. Everything else, a shell, eval, an interpreter, source, xargs, watch, ssh, cat, tee, an assignment, the command position, `\EOF`, `"EOF"`, `$'EOF'` and unquoted delimiters, text on the operator line, text after the terminator, an unterminated body and a word glued to `--body=`, keeps the old reading, the body lines read as commands. The legacy-reading cutoff is deleted with the machinery. shlex knows no ANSI-C quoting, so tokenize() keeps each `$'...'` string as one word (harmless `printf '%s' $'it\'s #\nprintenv\n'` passes; `bash -c $'printenv'` is read as `bash -c printenv`). Checked at this state: the oracle 65 cases, 0 mismatches; kw/guard_probe_final.py base candidate `loosened rows: 0`; the eight consumer variants and their rtk proxy and keyring exec forms 0 of 13 loosened; the reviewer's five here-document strings and the coordinator's eight consumer forms are BLOCKED rows and environment_dump; the four suites 151 tests, the host-copy comparison the only failure; 14 mutation controls of the recognizer, the consumer decision and the tokenizer change each make a test fail (two survivors found by the first run, a plain `<<` terminator with a leading tab and the second tokenizer reading, got a row each). The launcher-redirection strings and the corrections of the review's other findings follow in their own commits. Residual gaps, recorded as EXPECTED_PASS_THROUGH rows: a git or gh option that runs its value (`git rebase --exec "$(cat <<'EOF' ...)"`, `gh alias set`), an echo piped into a shell or written into a script that runs later, and an unquoted here-document that expands `$(...)` between single quotes (the base guard passed it too). Friction, in the strict direction: prose in a body that starts a line with `printenv` is refused outside the idiom, and so are `--body="$(...)"`, `cat` and `tee` as consumers and a text with an unbalanced apostrophe before the idiom. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The review of 172596e found that `env -u < "$PAPER_ENV_FILE" UNUSED cat` and `systemd-run --pipe --unit < "$PAPER_ENV_FILE" demo cat` are credential_file_read at the base guard and None at the tip: the index-based launcher walk took the redirection operator for the value of `-u` and of `--unit` and dropped the segment that held the credential-file operand. Reproduced with a matrix of redirection positions: for env, env -i, systemd-run, sudo, timeout, timeout -s, nice, rtk proxy, a keyring exec and four chains (sudo env, env sudo, nohup timeout), every position between the hops, each of `<`, `<<<`, `2>`, `>`, `>>` and `&>`, to a credential pointer, `.env` and `~/.aws/credentials`, 684 combinations: 8 rows looser than the base guard, all `<` or `<<<` to the pointer in an option-value slot (`env -u < ...`, `systemd-run ... --unit < ...`, the same behind sudo and env). Failed first (the new tests against the guard of the previous commit): 170 failures, one being the tolerated host-copy comparison: 152 subtests of the position matrix, 13 reviewer strings, and the wrapper tests behind rtk proxy (2) and a keyring exec (1) and the direct table (1). Design, following the coordinator's instruction: the raw segment is read as written beside every launcher walk (command_segments appends it before any walk), so a redirection is checked wherever a walk would put it; and no option-value slot swallows a redirection any more. skip_redirections() steps over a redirection operator and its target, and wrapper_options, env_command_start, rtk_command_start and prefix_end use it wherever they look for an option, a value or the started command (only an operator token starts one: a bare number may be a value, `nice -n 5 > out cmd`). An input redirection that was stepped over is appended to the command that gets it, as if it stood after it (launcher_chain now returns it), so `env -u UNUSED < .env cat` is read as `cat < .env`, and a keyring exec and a leading here-string do the same. Nothing is carried for a number before an operator: `sudo 2>/dev/null -u root printenv` stays a recorded gap. Checked at this state: for every launcher, position and operator the verdict equals the verdict of the same command with the redirection written last, 684 combinations, 0 mismatches; against the base guard 0 rows are looser and 166 are stricter, each the operand of a `cat` that the base guard only saw when the redirection stood after the command; kw/guard_probe_final.py and the oracle unchanged (`loosened rows: 0`, 65 cases, 0 mismatches); the smoke row of each reviewer string is credential_file_read; the four suites 153 tests, the host-copy comparison the only failure; 22 mutation controls (the earlier 14 and eight for the walkers, the raw segment and the carried redirections) each make a test fail. `nice > /tmp/out -n 5 cat .env`, a recorded gap of 172596e, is now a BLOCKED row. The redirection rows also cover a launcher with no command (`sudo < "$PAPER_ENV_FILE"`), which only the raw segment sees. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…oosenings and the ps friction The review's other findings, each a correction of a claim, not of a rule. systemd-run provenance. The guard's hint, docstring, comments and docs said that the manager logs systemd-run's command line, an `-E NAME=value` included, in the journal. Read from upstream (systemd v255, src/run/run.c, fetched 2026-09-29): the default Description, which the manager logs as `Started <unit> - <description>`, is `quote_command_line(arg_cmdline)` (lines 1940 to 1951), and arg_cmdline is the command and its arguments after the options; `-E` fills arg_environment (line 348), which is appended to the start message as the unit's Environment property (lines 853 to 866). So a value given with `-E`, `--setenv` or `-p Environment=` is in the command line of the systemd-run process, which the process listing shows while it runs, and in the transient unit's Environment property on the user bus (`systemctl --user show -p Environment UNIT`, and any bus client); it is not in the journal, while a value written after the command is (a recorded gap). The rule, a secret name on a systemd-run line is refused, is unchanged. The verbatim commit message of cf58426 that the tests carry as a real message keeps its old wording, as it is a historical text. Failed first: test_the_systemd_run_hint_says_where_a_value_on_the_command_line_goes failed against the previous guard (the hint named the journal); the rest of this commit is documentation and rows that the previous guard already satisfied (ALLOWED rows for `ps -fu steve`, `ps -fu eve` and `ps -u steve` pass there too). The list of loosenings. The base guard c26800f refuses `ps -fu steve` and `ps -fu eve`, because it reads the user name as a dashless BSD cluster with an `e` (`ps -u steve` it passes); `-u` takes a value, and this work lets them pass. Together with the canonical idiom of the here-document item, that is the only loosening against the base guard: docs/secret-storage.md now says so in one item and the ALLOWED rows carry it. `ps -CEmacs` stays refused, documented as friction (on procps `-C` takes a command name: `ps -C emacs`). Checked: the oracle 65 cases, 0 mismatches; kw/guard_probe_final.py `loosened rows: 0`; the four suites 154 tests, the host-copy comparison the only failure. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ected measurements My own adversarial probe of the exemption before the final pass, nested and wrapping constructs against the base guard, found that `eval $(echo "$(cat <<'EOF'` + `printenv` + `EOF` + `)")` passed (the base guard passed it too, so it is no regression, but bash runs `printenv` there): the tokenizer splits an unquoted `$(echo IDENT)` into a segment of its own whose command is echo, and exempt_idioms() judged the idiom by that segment, while the substitution's output goes to eval. The same held for `bash -c $(echo IDIOM)`, `bash <<< $(echo IDIOM)`, `x=$(echo IDIOM); eval $x`, `xargs sh -c $(...)`, `python3 -c $(...)`, `eval $(git log -1 --format=%B IDIOM)`, a subshell or brace group around any of them, and `echo $(git commit -m IDIOM)`. The dq-substitution rule of the previous commit covered only the substitutions inside double quotes. Failed first (the new tests against the guard of the previous commit): 68 failures and 16 errors, one failure the tolerated host-copy comparison: 15 rows of the exemption table, the 14 new BLOCKED rows directly (13), behind rtk proxy (26) and behind a keyring exec (13), and the two tests of the new scan result (10 and 6 errors). Design: scan_shell() now returns a fifth result, the (start, end) of the text inside each outermost command substitution, `$(...)` or a backquote pair, quoted or not (a counter of open substitution frames, kept right through the `$((printenv) )` conversion and an unterminated substitution); exempt_idioms() refuses any idiom whose marker lies inside one, which replaces the bodies-only rule. A subshell `( ... )` and a brace group are no substitution: `( cd sub && git commit -m IDIOM )` stays exempt. The probe's 28 rows: 0 looser than the base guard; every nested form above is now an environment_dump; `git commit -m IDIOM` and `echo IDIOM` pass; the pipe into a shell stays a recorded gap. Docs: the exemption item names the rule (`eval $(echo "$(cat <<'EOF' ... EOF)")`), and the measured numbers of the item are corrected to the final ones: the consumer matrix is 40 consumers x 18 launcher prefixes x 5 here-document forms x 8 payloads x 4 substitution forms = 115,200 commands (the item said 39 and 112,320), the commit history is 1,956 messages (13 refused by the base guard, 10 by the guard, all by raw-text rules, 4 that the base refused pass), the differential over 60,493 derived commands loosens 102 rows and every one is a `ps -fu steve` or `ps -fu eve`, and the mutation fuzz of 24,840 mutants of 414 blocked strings loosens 38 to 44 per seed, of which 36 to 41 are variants of those two rows and 2 or 3 are mutants whose broken terminator leaves a real quoted body (the item said 24,720 mutants, 412 strings and 2 or 3, without the ps rows). Checked: the oracle 65 cases, 0 mismatches; kw/guard_probe_final.py `loosened rows: 0`; the four suites 155 tests, the host-copy comparison the only failure; 22 mutation controls each make a test fail (the rule that this commit strengthens is killed by the new rows). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…her strings are table rows
Housekeeping the design change left: the guard's module docstring still described the superseded
rule ("a quoted here-document inside a substitution whose command is git, gh, echo, printf, cat or
tee"), and said backslash-newline is joined first without saying that the canonical idiom is looked
for before that. It now says: a here-document's lines are command lines like any other; the one
exemption is a strict canonical idiom, a double-quoted word of exactly the shape
`"$(cat <<'EOF'` newline, body, `EOF` newline, `)"`, behind git, gh, echo or printf, whose body is
data for the command reading while the raw-text rules still read it. Three test comments that
named "data consumers" or "the first repair round" follow.
The two launcher strings of the review, `env -u < "$PAPER_ENV_FILE" UNUSED cat` and `systemd-run
--pipe --unit < "$PAPER_ENV_FILE" demo cat`, were checked by a test that asserts only that they
are refused; the coordinator asked for them in the tables, so they are BLOCKED rows
(credential_file_read, as at the base guard), where the wrapper tests also run them behind rtk
proxy and a keyring exec.
No behaviour changes: the four suites and the oracle pass as before, and the guard's hash in
SHA256SUMS moves with the docstring.
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A grammar fuzz of my own, launcher chains with redirections at random places and idioms in random placements, generated against the base guard (10 seeds of 60,000 commands: 5 with the launcher walkers of the previous commits), found 4 commands in 120,000 that the base guard blocks and the guard allowed: a keyring exec with a redirection between the variable name and the `--`, for example `python3 scripts/kernel_keyring.py exec tavily_api_key TAVILY_API_KEY < x.txt -- set` (base: environment_dump_in_keyring_exec). The base guard dropped that redirection when it took the command after `--`; the walkers of this work carry a skipped input redirection to the command it belongs to, and is_environment_dump() counted the carried `< x.txt` as arguments of `set` (`set a b` sets positional parameters), so the dump was no dump. `set < FILE` and `set > FILE` still print every variable. Failed first (the new rows against the guard of the previous commit): 53 failures, one the tolerated host-copy comparison: 16 BLOCKED rows directly, 28 behind rtk proxy, 8 behind a keyring exec. Four of the new rows contain a keyring exec of their own and live in KEYRING_BLOCKED, which the wrapper tests do not prefix with another one. Design: the module docstring already says "redirection operands are never taken for arguments"; is_environment_dump() now reads set, export, declare and typeset through command_arguments(), which drops the redirections, so `set > /tmp/vars`, `set < /dev/null`, `export -p > FILE`, `declare -p > FILE`, `typeset -p 2> err` and the keyring forms are an environment_dump, while `set -e > /dev/null`, `export FOO=1 > /dev/null` and `declare -a items > /dev/null` (a real argument) pass. This only tightens: the base guard passed the redirected forms. Checked: 10 seeds of the grammar fuzz, 600,000 commands, 0 looser than the base guard (seeds 1 and 2, which found the four, included); the four suites and the oracle unchanged; the docs sentence is in the launcher-redirection item. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The here-document item of the guard section still quoted the numbers of an earlier commit of this series (1,956 commit messages, 13 and 10 refused, 4 passing, a differential over 60,493 commands, mutation fuzz of 24,840 mutants of 414 strings). Replaced with the numbers measured at 660e681, the tip whose guard is byte-identical to this one: 1,960 messages, 10 refused (base 14), 5 that the base guard refused now pass, a differential over 62,365 derived commands that loosens 102 (all `ps -fu steve` or `ps -fu eve`), a grammar fuzz of 10 seeds of 60,000 launcher chains that loosens none, and mutation fuzz of 25,140 mutants of 419 blocked strings that loosens 33 to 44 per seed (29 to 38 `ps` variants, 4 to 6 mutants whose broken terminator leaves a real quoted body that bash prints and does not run). The loosenings item says 5 messages at 660e681 instead of 4. Two over-long lines of the item re-wrapped. Docs only: the guard is unchanged (sha256 72d43a2a15c4ccc002a13d9747ed818ee51832084811753912d4e9aa4ebf347e, 94,067 bytes), so the SHA256SUMS line stays. The four suites: 155 tests, one failure, the tolerated host-copy comparison (the host profile copy is not reinstalled), 3 skipped. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…'s history The measurement of the here-document item said "1,960 commit messages reachable there" at 660e681. The script that counts them (git log --all) reads every ref of the shared repository, including other branches and worktrees, so the number is the distinct messages on all refs when measured (1,964 on the run that found this) and grows without any change to the guard. The item now says that, names the guard as the one of 660e681, and the loosenings item points at that measurement instead of repeating "at 660e681". Re-measured on the current refs: 10 messages refused by the guard, 14 by the base guard, 5 that the base guard refused now pass (3899d49, 2f097bf, 172596e, c3bdbaf, 5b96461), so the sentences hold. The timeout item claimed that both readings of a command always run inside the hook's 10 s at the 200,000-character limit. It now carries the measurement behind that: seven inputs of about 199,000 characters that make every reading run (double-quoted text with a `#` throughout plus one idiom or 1,000 idioms, the same single-quoted, ANSI-C strings, one quoted word, many backquotes, many subshells) took 0.2 to 3.3 s in check() on this host, the slowest being the double-quoted text plus one idiom. Three over-long lines (the loosenings item and two in the timeout item) re-wrapped. Docs only: the guard is unchanged (sha256 72d43a2a15c4ccc002a13d9747ed818ee51832084811753912d4e9aa4ebf347e, 94,067 bytes) and the SHA256SUMS line stays. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…read as command lines behind every consumer The second verification review (GPT-6 through the gateway lane, of 50ca6ca) found six more blocking regressions, five of them inside the strict canonical-idiom exemption that replaced the general here-document reading: an idiom head inside another here-document's body that swallowed its terminator, an idiom inside an unquoted here-document that a shell reads, an echo piped into a shell, `git rebase --exec` (a git option that runs its value), a case pattern's `)` that closed a substitution early, and a process substitution that runs a shell. Each is an environment_dump at the base guard and None at that tip. The coordinator's decision: the exemption is a net risk for its friction benefit, so it goes, and the guard is tightening-only again. Two reviews and twelve block-to-allow regressions show that telling a data consumer from a code consumer by the words of one command line is what fails. Failed first (the six reviewer strings as BLOCKED rows against 50ca6ca): 24 failures, 6 direct, 12 behind rtk proxy and 6 behind a keyring exec; the two `ps -Ccat e` strings of the same review are the next commit. Design: idiom_spans, exempt_idioms, receiving_command, neutral_reading and everything only they used (IDIOM_HEAD, IDIOM_TAIL, IDIOM_FOLLOWERS, IDIOM_CONSUMERS, RESERVED_STARTERS, SUBSTITUTION_MARKER, MAX_LAUNCH_HOPS, the bisect import, and scan_shell's fifth result with its open-substitution counter) are deleted. check() reads the joined command as it did before the idiom existed: a here-document's lines are command lines, quoted delimiter or not, behind git, gh, echo and printf as behind a shell, eval, an interpreter, source, xargs, watch and ssh. The raw-text rules, the double-quoted substitution reading, comments, the launcher redirection handling, the systemd-run family, ps -E, systemctl and CANARY_E2E_KEY stay. Tests: the six strings and 16 documented-friction rows are BLOCKED rows (a commit message or pull-request body whose prose has a line that reads as a dump; a literal `$(printenv)` in a quoted here-document in file-writing, Python, commit-message and PR-comment workflows; Python's `set()` after a comment line in an interpreter's here-document), ALLOWED has the workarounds (`git commit -F`, `gh ... --body-file`) and a prose message that passes; the five idiom tests and their tables are gone, the real commit messages of this repository are recorded as the friction they now are (each refused with its reason), and a matrix of 26 consumers, 7 delimiter forms and 10 dumps (1,820 inert strings) must all be refused, so an exemption that returns for one consumer fails there. Checked: oracle phase base 65/65; the probe with the base guard `loosened rows: 0`; the four suites 151 tests, the one failure the tolerated host-copy comparison; every reviewer string of both reviews is blocked. Of 2,015 distinct commit messages on all refs, in the `git commit -m` and `gh pr create --body` patterns, the guard refuses 36 and the base guard 16, and none that the base guard refuses passes. The seven worst-case shapes of about 199,000 characters now take at most 1.0 s (3.3 s with the third reading). The guard section of docs/secret-storage.md states the friction and the workaround, and its loosenings item says one; SHA256SUMS carries the new guard hash. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…two readings and either one refuses The second verification review of 50ca6ca found that removing `-C` from the value-taking options of a cluster lost the procps reading: `ps -Ccat e` and `ps -fCcat e` (environment_dump at the base guard) returned None. On procps `-C cmdlist` takes a command name (`cat`) and the BSD `e` after it shows the environment; on macOS `-C` is a flag, so `-Ccat` is `-C -c -a -t` and `t` takes `e` as a tty. The coordinator's instruction: keep both readings and refuse when either shows the environment. Failed first (the new rows against 752def7f): 16 failures, 4 direct (`ps -Ccat e`, `-fCcat e`, `-Ccat eww`, `-Ccat E`), 8 behind rtk proxy and 4 behind a keyring exec. `ps -Cnginx e` and `-fCnginx auxe` were refused before by chance (the `g` of `nginx` is a value letter) and stay refused. Design: ps_reading_shows_environment(words, procps) is one host's reading, ps_shows_environment() the union. procps (PS_ARG_OPTIONS, `-C` a value option, dashless BSD words anywhere, no dashed `-E`) and macOS (PS_MACOS_ARG_OPTIONS, `-C` a flag, dashed `-E`, a dashless option string only as the first argument) each keep their own state of which word a value option takes. The `-C -E` special case of the hybrid reading is gone: the macOS reading finds it. Sources: procps-ng 4.0.4 ps(1) on this host (`-C cmdlist`, `e` "Show the environment after the command", no `-E`), and Apple adv_cmds ps/ps.c at 60bc9ebf (fetched 2026-09-29): PS_ARGS `aACcdeEfg:G:hjLlMmO:o:p:rSTt:U:u:vwx`, `case 'C': rawcpu = 1`, `kludge_oldps_options(..., argv[1], ...)` applied to argv[1] only, and a later non-option word is a process id or an "illegal argument". Checked on this host: `ps -Cbash u` honours the BSD `u` after the glued command name and `ps -fC bash u` reports conflicting format options, so a dashless word after a UNIX option is a BSD option on procps. `ps -CEmacs` stays refused (friction: the macOS reading; write `ps -C emacs`). Loosening, the one this work makes, now stated generally: a value after a clustered value-taking option is no BSD flag cluster (`ps -fu steve`, `-fo user`, `-ft e`, `-fU steve`, `-fC e`); the base guard skipped the value after a stand-alone `-u` but not after `-fu`. A differential over 478,915 ps command lines (every one of up to three words from a vocabulary of 65, plus 200,000 random ones of four to six words) loosens 6,673 against the base guard, every one holding such a cluster with its value (0 outside that family); the reference two-host model of the tests (a per-letter state machine) and the guard agree on all 478,915. The tests carry that reference over 160,434 lines (a vocabulary of 54), ALLOWED rows for the value forms and BLOCKED rows for the union. Checked: oracle 65/65; probe with the base guard `loosened rows: 0`; the four suites 152 tests, the one failure the tolerated host-copy comparison. The docs replace the ps item and the loosenings item, and record macOS's legacy mode (`-e` read as `-E` when ps.c's `u03` is off) as a gap. SHA256SUMS carries the new guard hash. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… refused as command_too_complex The second verification review found two commands far below the 200,000-character cap that outlast the hook's 10 s timeout, where a timed-out PreToolUse hook blocks nothing: a keyring exec with 1,200 `python3` words (9,645 characters; the analysis read every suffix of them as a command, 5.8 million characters through shlex: base 2.1 to 2.8 s, 50ca6ca 12 to 13 s) and 195,000 emoji in five nested double-quoted substitutions (195,068 characters; four-byte characters cost four times as much to tokenize and the text was read at every level: base 1.6 s, 50ca6ca 15.7 to 17.9 s here, over 25 s in the review). My own scan found the same class elsewhere: 4,000 distinct keyring execs (25 s), 3,500 keyring execs of one variable, each starting a distinct program (268 s: a quadratic test of mentions against spans and a scan of the whole text for each started command), a chain of 6,000 nested keyring execs (a copy of the rest of the words at every hop, past 30 s), and a keyring exec whose started command is a 180,000-character argument (6.4 s). Failed first (the new tests against d46f6c0c): the six unit tests 2 failures and 10 errors (no WorkBudgetExceeded, WORK_LIMITS, storage_width or start_work, expand kept 3 copies of one segment, no command_too_complex hint); the timing rows in the tests' own child process: 1,200 words 12.48 s, 600 words 3.16 s, 5,000 words past the 45 s limit, the emoji input 17.20 s, 120,000 euro signs 11.0 s. Design: check() starts a budget, reads the command, stops it (start_work, stop_work; outside a check call nothing is counted). WORK_LIMITS: (a) characters passed to shlex over every reading and nesting level, each at its storage width (1 byte ASCII and Latin-1, 2 for the rest of the BMP, 4 beyond: shlex costs 0.41, 0.85 and 1.70 s on 200,000 characters of one quoted word), 400,000; (b) texts read 10,000, words emitted (each segment its words plus one) 1,000,000, launched-command reads of the keyring analysis 500. lex() charges before shlex runs, command_segments() charges each text and each emitted segment, launched_commands() and keyring_reason() charge each read. Over a limit check() raises WorkBudgetExceeded (args[0] names the counter), never a reason, so no caller can take an unread command for an allowed one; main() prints `blocked (command_too_complex)` with one line, no command text, and the hint to split it or use the Write tool, and exits 2. Algorithmic fixes that need no budget: expand() and the keyring analysis drop identical segments, keyring_reason() scans the text once for each pattern (a cache), and mentions_injected_variable() tests a mention against the starts of the exec arguments (a set) instead of every span. The dead `strict` parameter of lex() and tokenize() went with the idiom code. Limits are from measurement, not a guess: each counter alone holds its worst case near a second (398,000 characters 1.1 s, 9,900 texts 0.3 s, 1,000,000 words 0.2 s, 490 reads 0.03 s); the largest real command of this repository (an 82,000-character script through a here-document) spends 35% of the characters and 1% of the words, the largest commit message 57%, none of 785 fenced blocks, 4,338 fenced lines, 201 scripts or 1,877 commit messages comes near a limit; 500 random mixes of 22 adversarial building blocks at 199,000 characters took at most 0.9 s, and 64 shape families at 25,000 to 199,000 characters at most 1.1 s. The reviewer's two inputs are refused in 0.15 s and 0.04 s, their 600- and 5,000-word variants in 0.15 s. Tests: 33 timing rows with the 3 s bound (the reviewer's two, 600, 60 and 5,000 words, euro signs, nested and distinct keyring execs, a chain of 6,000, the launched programs and the 100,000-character argument, each with the deterministic verdict, `command_too_complex` for what the budget refuses), counters tripped one at a time and the same shapes just inside, the storage width, the accounting of one expand() call, the deduplication, main() and the real hook on both inputs (under 3 s), and check() may raise only WorkBudgetExceeded on a pathological row. Rows of the earlier rounds that the budget now refuses were resized (a 160,000-character 20,000-deep nesting is refused, 5,000 deep reports its dump). Checked: oracle 65/65; probe with the base guard `loosened rows: 0` (its 1 MB quoted word now refused in 0.15 s, was 10.4 s); the four suites 157 tests, the one failure the tolerated host-copy comparison. The guard section states the counters, the numbers and the friction; SHA256SUMS carries the new guard hash. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… from `-E NAME` The second verification review's note: the provenance still conflated an explicitly supplied value with value-less environment forwarding, in docs/secret-storage.md and in the module (the hint and the docstrings). Read from systemd v255 (fetched 2026-09-29): `-E NAME=value` puts the value in the systemd-run process's own argv (the process listing shows it while systemd-run runs) and in the transient unit's Environment property; `-E NAME` puts only the name in argv, and systemd-run takes the caller's own value from its environment (`strv_env_replace_strdup_passthrough`, src/basic/env-util.c:417, called for `case 'E'`, src/run/run.c:348), so that value reaches the Environment property and no argv. The property is appended to the start message over the user bus (`arg_environment`, run.c:853-866), which `systemctl --user show -p Environment UNIT` and any bus client read. Not the journal: the unit's description, which the manager logs as `Started <unit> - <description>`, is the started command and its arguments after the options (`quote_command_line(arg_cmdline)`, run.c:1940-1951), so neither `-E` form is in it; a value written after the command is (a recorded gap, unchanged). Failed first: test_the_systemd_run_hint_says_where_a_value_on_the_command_line_goes against the previous hint (which named the process listing and the user bus for every form): 1 failure. Changed: the hint of `secret_variable_on_command_line` (it now says `-E NAME=value` puts the value in the command line, `-E NAME` forwards the caller's value, and either way it lands in the Environment property on the user bus, readable with `systemctl --user show -p Environment UNIT`; still no word about the journal), the docstring of systemd_run_sets_secret, the module docstring and the docs paragraph of the systemd launchers, with the three source lines cited. The rule itself is unchanged: the same names on the same forms are refused (the review's `systemd-run --user -E GH_TOKEN true` is a row of the hint test). Checked: oracle 65/65; probe with the base guard `loosened rows: 0`; the four suites 157 tests, the one failure the tolerated host-copy comparison. SHA256SUMS carries the new guard hash. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…gacy `-e` as a recorded gap, a mutation the word limit survived Tests only; the guard is unchanged (sha256 678cd174f3c8bfbe936a8485e21902ed82c2580fd5cc21c4a6b4bbf0a21a9b29, 93,563 bytes), so the SHA256SUMS line stays. - BLOCKED: `ps` followed by 70,000 `E`, a letter that is no flag and a dump, the first review's input in full (the timing row PATHOLOGICAL already bounded its time; the BLOCKED table also runs it behind rtk proxy and a keyring exec). - EXPECTED_PASS_THROUGH: `ps -e`. macOS's ps.c falls `case 'e'` through to `case 'E'` (the environment display) when `u03`, its unix2003 compatibility flag, is off; the guard reads a dashed `-e` as every process, as procps and macOS's default do, since refusing it would refuse every `ps -ef`. The row is the recorded gap of the ps item of docs/secret-storage.md. - The work-budget test gets a second `words` case: 40 keyring execs in front of a 60,000-word tail emit 2.4 million words at the hops. A mutation control of this round (the word limit ten times larger) survived the first version, since the 6,000-hop chain still raised 'words' at 10 million within the 3 s bound; the new case does not raise at 10 million (the character budget trips later, in the keyring analysis), so the mutant is killed. Mutation controls of this round: 24 mutants of the guard (a here-document body becoming data again; the ps union cut to either reading, `-C` a value option on macOS, a dashless cluster at any position on macOS, a cluster that takes no value, a dashed `-E` on procps; the storage width, each of the four counters and each of their limits, the mention scan charge, deduplication, main() not refusing a spent budget, the budget left running, the mention test counting exec arguments, a budget that returns instead of raising; the systemd-run hint's value-less form), each run against the whole test module in a scratch export with an 8 s limit per timing row: 24 killed, 0 survived after this commit (23 and the one above before it). The docs sentence about the work budget no longer names a scratch measurement tool (one word). Checked: the test module 41 tests, the one failure the tolerated host-copy comparison. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…unt named by its guard Docs only; the guard is unchanged (sha256 678cd174f3c8bfbe936a8485e21902ed82c2580fd5cc21c4a6b4bbf0a21a9b29, 93,563 bytes), so the SHA256SUMS line stays. - The loosenings item carries the final differential measured with the guard of 36c847db (the tests changed after it, not the guard): 66,176 commands derived from 1,079 table and oracle rows loosen 510, which are 10 `ps` forms (`ps -fu steve`, `-fu eve`, `-fo user`, `-ft e`, `-fO user`, `-fU steve`, `-fG eve`, `-fg steve`, `-fC e`, `-fC eww`) in 51 wrappers each (8 launcher prefixes and 6 suffixes, `bash -c`, `sh -c`, `eval`); three seeds of the mutation fuzz (26,400 mutants of the 440 strings the base guard blocks, per seed) loosen 153, 175 and 173, and the guard's own segment reading puts every one of those 501 (and the 510) in the clustered-value `ps` family, 1,011 of 1,011; nothing is loosened by the 785 fenced blocks, 4,338 fenced lines, 201 scripts twice, 16,119 distinct repository lines, two fuzzers of 60,000 random strings, ten grammar fuzzes of 60,000 launcher chains with redirections, the 40-consumer matrix of 115,200 commands and the 684 launcher redirection positions. A differential over 478,915 `ps` command lines (every one of up to three words from a vocabulary of 65, and 200,000 random ones) loosens 6,673, none outside that family. - The timeout item: 1,000 random mixes of the 22 adversarial building blocks (four seeds) took at most 0.91 s, and the figures are runs on a shared host (the same 192,000-character quoted word took 0.8 s alone and 1.7 s at a load average of 8), so a figure reads within a factor of two. - The here-document item names the guard its commit-message count was measured with (752def7f, 2,015 messages then); the final run of this round, 2,065 messages, gives 41 refused by this guard (this round's own messages, which name printenv and env in prose, are among the new ones) and 16 by the base guard, and none that the base guard refuses passes. Checked: the four suites 157 tests, the one failure the tolerated host-copy comparison; the whole repository discovery 7,161 tests, 2 failures (that one and the publication validator, which reports the pin mismatches of the four changed files, the two files that changed in 172596e without a re-pin, and nothing else). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The third verification review found that the clustered value-taking ps loosening lets procps personalities through: with PS_PERSONALITY=old or I_WANT_A_BROKEN_PS (on the command line or inherited from the shell, which a hook cannot see) procps parses `ps -axu e` BSD-style, where `u` takes no value and `e` shows the environment. The coordinator's decision: remove the loosening entirely. ps_shows_environment now also applies the c26800f rule verbatim (prior_ps_shows_environment, PS_BSD_CLUSTER): a dashless word of that guard's cluster alphabet with an `e` is refused wherever it stands, unless a stand-alone value option takes it. The procps and macOS readings stay for everything else. Tests: the review's three strings and the brief's rows move to BLOCKED (environment_dump); the ALLOWED rows that recorded the loosening are gone; the ps state-machine test adds an independent reference for the c26800f reading. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…er reading refuses The third verification review and the coordinator's probe found commands that the base guard (c26800f) refused and this version let through: systemd-run's --description took the keyring script as its value, so the keyring exec it names was never unwrapped. A probe of the same class found more: the option values of run0, systemd-cat and systemd-inhibit, a path-qualified wrapper whose option value is the script, a systemd-run inside a keyring exec, and a redirection operator that the base walk took for the value of `sudo -u`, `nice -n` or `env -u`. Instead of finding each walk that differs, check() now reads the words twice: this version's reading first, and when it allows the command, the reading of c26800f (prior_reading: that guard's walk and rules, ported function by function, spending from the work budget, with `find -exec` walked by index and identical started commands read once). A command either reading refuses is refused, so no command the base refused passes. lex() keeps each text's words for the call, so the second reading does not tokenize the command again. Measured in-process against the base guard file (sha256 f9be81b2): over 443,463 commands (test tables in 17 launcher prefixes and 6 suffixes, fenced lines, 19,296 launcher option-value rows, 10 seeded mutants of each of 33,502 base-refused strings) the prior reading equals the base on every one and check() loosens none. Long env or rtk chains that end in a harmless command now spend the words budget in the prior reading and are refused as command_too_complex in under 0.2 s (the base took 29 s and 56 s). Tests: the four probe strings and eleven more of the class as BLOCKED rows with the base's reasons; a test that each passes this version's reading and is refused by the prior one; PATHOLOGICAL rows for the harmless-ending chains. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…; plain text skips shlex
The third verification review found raw STORE_PATHS regular expressions that
run before any work accounting and read an unbounded run again from every
repeat of the literal before it: 102,016 characters of XDG_CONFIG_HOME:- took
11.9 s, HF_HOME:- 13 to 14 s and /proc 13 to 15 s, past the 10 s hook timeout
that fails open. A probe of every text rule found a fourth, the `gh auth
status ... -t` pattern.
Each of the four is now a LinearScan whose search() answers as the old
regular expression: it splits the text once into the runs the pattern cannot
cross (a ${NAME:-default} run of neither `}` nor blank, a path run without a
blank or `//`) and searches each run for literals; the gh rule compares sorted
positions. Old against new: 7,229,043 texts a pattern (2,000,000 random and
every text of up to six fragments), 0 differences; the unit test keeps 100,000
random texts a pattern and every text of up to four fragments.
With the regular expressions linear, one unquoted word of 199,000 characters
still cost 0.53 s in shlex, which copies the word at each character. A text
with no quote and no backslash is now split by SHLEX_PLAIN, which shlex's state
machine reduces to there (its four blanks end a word, a run of its punctuation
is a word, a `#` starts a comment in the legacy reading): 2,111,111 texts in
both readings, 0 differences against shlex.
Measured in-process: the review's three inputs take 0.07 to 0.18 s, and every
STORE_PATHS leading literal repeated to 199,000 characters at most 0.24 s,
before a dump or a harmless command. Tests: those timings (0.5 s bound, the
real hook under 1 s on the review's inputs), both equivalences, and the
backtracking test now repeats each pattern's leading literal as well.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The third verification review measured ('find . ' + '-ok ' * 49000 + '; ...')
at 3.9 s: reader_arguments sliced the rest of the words for every action to
strip its prefix. A prefix that runs to the end from every action
('-exec sudo' repeated, 18,000 times) took more than 20 s however it was
sliced, since every action walked it again.
prefix_end and the prior reading's prior_prefix_end take an optional `ends`
map: a walk from a position ends where the walk from its next step ends, so
the answer is kept for every position a walk steps on and each position is
walked once. Both find readers (this version's and the prior reading's) use it.
Measured in-process: the review's input 0.23 s, its harmless-ending twin (which
reaches the prior reading) 0.32 s, the -exec sudo chain 0.43 s at 330,013
characters. Tests: the four find shapes join the review's timing rows (0.5 s
here, wider in CI; the real hook within a second).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…operator
The third verification review found `set 1 > out` and `set 3 < input` newly
refused (the base guard passed them): the dump rule of set, export, declare and
typeset dropped every redirection from the arguments, and read a number before
an operator as its descriptor because shlex splits `1>out` and `1 > out` alike.
In bash the number is the descriptor only when it touches the operator; apart,
it is an argument (here, a positional parameter).
lex() now marks each number (or bash's `{name}`) that touches `<` or `>`
before it splits the text and hands it back as a Descriptor, a str that equals
the number, so every other rule reads it as before; the dump rule counts only a
Descriptor as a redirection (command_arguments(..., touching_only=True)). A
quoted or glued number is no descriptor, and a mark inside a quoted word is
removed. A text that holds the mark character itself is not marked, and every
number before an operator counts there, as before.
Verdicts against the base guard: `set 1 > out` and `set 3 < input` pass (base:
pass); `set 1>out`, `set >out` and `set 2>/dev/null` print every variable and
stay refused as environment_dump (base: pass).
Tests: the two review rows and four more in ALLOWED, five touching forms in
BLOCKED with the base verdict noted, and a lex() test of the marks against
shlex's words.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e-pass claim is corrected Under a load average near 20 on this shared host, the harmless-ending find shape of the last commit took 0.54 s of wall clock against the 0.5 s bound. A profile showed three Python-level loops that do not depend on the input's risk: the default-run scans (XDG_CONFIG_HOME, HF_HOME) and the /proc scan walked every run of the text (49,000 there), each find action walked its prefix even when the next word is an option, and is_environment_dump built the argument list of every segment. Now: - default_run_scan and proc_environ_scan read only the runs that hold their literal, from each first occurrence to the end of its run (the same answers: 7,229,043 texts a pattern against the c26800f regular expressions, 0 differences); - a find action followed by an option word (dash-led, no slash) is skipped in both readings: the walk would stop on that word, which is no reader; - is_environment_dump reads the arguments of set, export, declare and typeset only. Measured in-process: the harmless-ending find shape 0.11 s (was 0.32 s), every STORE_PATHS leading literal at 199,000 characters at most 0.12 s. The timing test bounds processor time at 0.5 s (wall clock 3 s), since the host is shared. The module docstring no longer says that every scan reads a text once (the third review showed four STORE_PATHS patterns that did not); it names the LinearScan rules, the plain split, the descriptor mark and the two readings. The WORK_LIMITS comments say both readings spend from one budget and give the 2026-09-30 measurements. The oracle table comment counts the added rows as the review asked: 245 of the 342 added BLOCKED and KEYRING_BLOCKED rows were allowed by the base guard, 97 are regression controls. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…iptor adjacency, recount docs/secret-storage.md, guard section, dated 2026-09-30 statements: - The opening paragraph and the former "Loosenings" item: the clustered-value ps loosening and the claim that it was the only block-to-allow difference are withdrawn; the new "No loosening" item names the third review's launcher strings and the class the probe found, the prior reading that closes it by construction, the evidence (the prior reading equals the base guard on 460,910 commands with the final guard; no harness loosens any command), the 44 reason changes that 6c4f63d already had, and the recount of added rows (245 new refusals, 97 regression controls of 342). - The ps item: personalities, the base guard's reading kept, the forms that are refused again, and agent-safe forms checked against this guard (`ps -f -U 1000`, `ps -fU 1000`, `pgrep -a -u steve node`). - The redirection item: a number is a descriptor for the dump rule only when it touches its operator (`set 1 > out` passes, `set 1>out` stays refused). - The timeout item: the one-pass claim is corrected (four STORE_PATHS patterns were quadratic), with the linear scans, the plain split, the find walk, the prior reading's chain limits and the re-measured budget and friction figures (sh -c strings 4,999; two commit messages of the refs over the characters budget, as with 6c4f63d). The guard's WORK_LIMITS comment gives the mix timing as measured over several runs (a comment-only change: the AST equals the guard the harnesses ran on); SHA256SUMS carries its hash. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ot-file commit, last) Hashes of the six changed pinned files on top of main's manifest at b4056a3 (docs/lanes.md protocol: main's copy, register each file, component_matrix and new_host_grand_list --write), plus the receipt guard-k3-verification-20260930 (artifact measurement: the acceptance oracle, probes, the documented-command replay, real hook timings and the two GPT-6 verification rounds of this change); validate.py passes: 69 components, 8297 hashed files, 175 receipts. The installed host copy is reinstalled after merge and after Gate A window W, announced first. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seathatflowsinourveins
force-pushed
the
claude/guard-launchers-20260929
branch
from
October 1, 2026 05:50
fb4a586 to
7dc5cf1
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Scope
What this PR changes, in one or two sentences: tightens the host-wide Claude Code command guard (
scripts/hooks/secret_path_guard.py). Changes:systemd-run,run0,systemd-inhibitandsystemd-catare read as launchers. Before this change a reader they start was not inspected:systemd-run --user --pipe --wait cat "$PAPER_ENV_FILE"returned the pointed-to file.$(...)and backticks inside double quotes are read as commands.ps -Eforms andsystemctl show-environment(and a baresystemctl show) are refused.check()call bounds tokenisation and launcher analysis (command_too_complex).CANARY_E2E_KEYis a secret name.There is no loosening.
Base commit:
6bbef3c6(main at the rebase; headdc33b48a)Lane:
lane:foundationOwned paths touched:
scripts/hooks/secret_path_guard.py,tests/test_secret_path_guard.py,adoption/hooks/claude/SHA256SUMS(the guard's pin line) and the guard section ofdocs/secret-storage.md. Three doc notes recommended a command the guard now refuses:adoption/templates/systemd/omniroute.serviceline 26,adoption/templates/systemd/token-report-refresh.serviceandtools/token-report/README.md. The change there is additive: the operator's command stays, with the agent-safe alternative added. Also the receiptevidence/receipts/guard-k3-verification-20260930.jsonand the last commit'smanifests/evidence.json. Not touched: the host copy (reinstalled after merge and after the Gate A window closes, announced first), any credential file.Not in this PR, stated plainly: the next guard change (K4) carries the following, from a written contract with its own review:
set(a) | set(b)in apython3 -body);/api/on the gateway ports, with an allowlist; decision record "OmniRoute management API (decision, 2026-09-30)");CLAUDE_CODE_OAUTH_TOKEN;SOTA sources
systemd/systemd@db11bab3):src/run/run.candsrc/systemctl/systemctl.coption tables: value options and boolean flags;-E NAME[=VALUE],-p Environment=.Descriptionis built from the started command's arguments after the options (run.c:1940-1951).-E NAME=valueputs the value in the caller's argv and in the transient unit's Environment property;-E NAMEforwards the caller's value into the property (env-util.c:417,run.c:348,run.c:853-866).adv_cmdsps/ps.cat60bc9ebf(PS_ARGS,case 'C',kludge_oldps_optionsapplied toargv[1]only) and ps(1).-Ctakes a command name; a dashless BSD word is read anywhere;PS_PERSONALITYandI_WANT_A_BROKEN_PSselect BSD parsing).psis refused when either host's reading shows the environment.redocumentation (the module backtracks; there is no linear-time guarantee): the four affected patterns were rewritten as explicit scans, with an old-pattern-versus-scan equivalence test.Evidence-class table
local_integrationguard-k3-verification-20260930.local_integrationguard_oracle.py: 65 cases, 0 mismatcheslocal_integrationlocal_integrationlocal_integrationlocal_integrationguard-k3-verification-20260930.source_reviewplus probes by the reviewergpt-6-astra, effort max, OmniRoute gateway lane), one before and one after the repair round (below)Local commands run
Review and repair record
git commitidiom exemption. Each was rejected after an independent GPT-6 review found six blocking block-to-allow regressions in it. Both are removed.b00ac4e1):PS_PERSONALITY=old ps -axu e) passed the clustered-valuepsloosening;find -okprefix walk (3.9 s);set 1 > outnewly refused;systemd-run --description kernel_keyring.py exec ...block-to-allow rows. The repair found 11 more of the same class.psloosening is removed.findprefixes walked by index; descriptor adjacency; docs corrected.set 0 < /dev/null; set 0</dev/nullpasses. The installed guard allows it too, so this PR does not loosen anything.\x01byte."$(cat <<'EOF' ... EOF)"whose prose has a line that reads as an environment dump is refused. Write it with the editor tool and passgit commit -F <file>orgh ... --body-file <file>.ps -fu steveandps -fo userare refused again, as by the installed guard.ps -f -U 1000andpgrep -u stevepass.envhops, or 997 two-word hops, iscommand_too_complex.Decision record
The guard section of
docs/secret-storage.mdanddocs/decisions/2026-09-29-key-management.md: the status section's item 1 and the dated 2026-09-30 sections.Host evidence
Not applicable to this PR: no file under
evidence/hosts/changes.Merge and reinstall happen only after the Gate A owner (native-agent-stack-2d) says window W is closed: the frozen check
hooks.carriers_match_repocompares the installed hooks with main. Then, announced first:gh pr merge 511 --squash --match-head-commit <head>.cp -p ~/.claude/hooks/secret_path_guard.py ~/.claude/hooks/secret_path_guard.py.bak-<date>.python3 tools/adoption/install_claude_profile.py --only guardfrom the live clone.tests.test_secret_path_guard.Checklist
permissions: contents: read: no workflow changes.🤖 Generated with Claude Code