Repository navigation
Integrate native tool evidence with explicit source-host scope - #7
Merged
seathatflowsinourveins merged 5 commits intoSep 20, 2026
Merged
Conversation
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 23, 2026
…okenizer, exact hostname) (#104) * Fix CodeQL first-analysis alerts: href scheme allowlist, case-insensitive tag scan, exact hostname check Resolves the 5 real defects from the repository's first CodeQL default-setup analysis (commit 168a3a8, alerts 1/3/4/5/8), re-located at base 796f759 since PR #96 changed the generated explorer between the two: - js/xss-through-dom (docs/ecosystem/template.html:145): link() now builds href through a safeHref() helper that only allows http:/https: URLs, blocking a javascript:-URI href from catalog data. - py/bad-tag-filter (scripts/build_ecosystem.py:704, tests/test_ecosystem_manifest.py:231, tests/test_claude_repository_evidence.py:147): add re.IGNORECASE so an injected uppercase <SCRIPT> tag is still counted into the page's inline-script CSP hash instead of silently bypassing the single-script precondition. - py/incomplete-url-substring-sanitization (tests/test_lifecycle_capture.py:103): replace the "sec.gov" in url substring check with an exact urlparse(url).hostname comparison. The remaining 5 alerts (2, 6, 7, 9, 10) are false positives / test-only synthetic-secret fixtures; their justification is recorded for the coordinator to apply via the GitHub API in docs/decisions/2026-09-22-codeql-first-analysis.md and the sibling codeql-dismissals.json, not dismissed here. manifests/evidence.json's sha256/bytes entries for the five touched files are refreshed so scripts/validate.py (run by validate.yml) still passes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Correct alert 7/9 location prose and add an executed safeHref check (review fixes) Resolves the independent reviewer's three medium findings on 65f73e2: alerts #7 and #9's location descriptions and dismissal comments pointed at the wrong code after the 796f759 re-location (both now cite the actual CodeQL sink); alert #1's "node -e smoke check" claim named no runnable command, so it is replaced with an executed unittest that runs the committed safeHref helper (extracted verbatim from template.html) under Node and asserts javascript:/data:/vbscript:/file:/mailto: are rejected while http(s) and relative URLs pass through, and alert #1 is relabeled defense-in-depth given build_ecosystem.py's public_url() is the primary control. Refreshes manifests/evidence.json for the two touched files so validate.py stays green. Re-running the full suite three times confirms the reviewer's flagged skip-count (338) is deterministic, not a flake. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Replace the <script> regexes with stdlib HTML tokenizers (CodeQL py/bad-tag-filter) re.I alone would likely leave py/bad-tag-filter open (it also flags missed end-tag variants such as </script >). The build script and both tests now use html.parser, which tokenizes like a browser; on the real generated pages the parsers return the identical single body the regexes returned. Corrects the record's alert #1 control description (loopback_url also admits http loopback links) and the skip-count claim. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * Fail loudly where html.parser and a browser disagree on <script> Review follow-up: a self-closing <script/> or a script after <!--> (Python 3.12) was invisible to InlineScripts. render_from_data now requires no self-closing script and requires the parser's script-start count to equal the raw "<script" count; the CSP test asserts the same count. The record's loopback_url and browser-equivalence wording is corrected. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: seathatflowsinourveins <234074349+seathatflowsinourveins@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
pushed a commit
that referenced
this pull request
Sep 24, 2026
… found
Both independent reviews confirmed round 3 mostly holds (all 24 earlier
mutations still killed) but independently found the same new regression:
- HIGH: ambiguity erased a confirmed action. AlpacaCorporateActionsSource.
fetch() used to overwrite a symbol's whole record list with
LOOKUP_AMBIGUOUS whenever ANY row for that symbol lacked a governing
date, discarding valid records collected in the SAME response (e.g. a
confirmed split plus an undated dividend for the same symbol).
CorporateActionMonitor._apply_success also overwrote a previously
confirmed record with a later per-symbol LOOKUP_AMBIGUOUS/LOOKUP_FAILED
result. Fixed at both layers: fetch() only ever falls back to the
sentinel for a symbol with ZERO valid records of its own in that
response; _apply_success degrades (keeping the last good record)
instead of overwriting whenever the new per-symbol result is a
sentinel.
- MEDIUM: raw rows with no identifiable symbol used to silently count as
clear. A row (mapped or unmapped bucket) naming no symbol at all now
marks every requested symbol ambiguous.
- MEDIUM: uncertainty discovered after the final rebalance() tick could
still end held_overnight. The outcome's corporate_action_guard summary
is now recomputed fresh (_final_corporate_action_guard_summary) against
the guard's current evaluate() and the truly final held positions,
immediately before the outcome is built, instead of trusting the
strategy's last-tick snapshot.
- LOW: shutdown errors could bypass refresh-task cleanup -- session.stop()/
port.stop()/the task wait are now in their own nested try, with
_shutdown_corporate_action_tasks in its own guaranteed finally.
- LOW: the private _get_marketdata call is replaced with the public
client.get() and this wrapper's own explicit next_page_token pagination
(matching blueprints/us-equities/alpaca-historical's own bridge pattern
for this endpoint), with region=us and a page-count ceiling that fails
closed if a continuation token survives it.
- LOW: one unmapped-type row used to fail the fetch for every requested
symbol. An unmapped bucket now degrades only the symbol(s) its own rows
actually name (Alpaca API reference, corporateactions-1, fetched
2026-09-24: partial_call/reorganization/capital_gains_distribution
documented as additional types; data_quality accepts complete/all).
- LOW (measured): a hung refresh thread used to delay run_native's own
result-save/account-lock release by the hang's full duration, because
asyncio.run()'s cleanup joins every asyncio.to_thread/default-executor
thread. The worker now runs on a daemon thread bridged back via
loop.call_soon_threadsafe; asyncio.run() never waits for it (measured:
a 4s hang delayed exit by ~4s before, ~0.1s after).
- NIT: a timeout processed immediately after a successful commit for the
same generation is now ignored (a new _committed_generation check),
instead of degrading an already-fresh result for the retry interval.
- Wording: corrected the "never overlapped"/"bounded fetch" claims to
state exactly what actually bounds each thing, and the native-check
record's "2026 is the only calendar year" claim (2027 is also covered).
Adds targeted tests for every finding above, including tests that go
through the real HTTP-mocked SDK pagination path (a two-page response
with a real next_page_token, and the cap enforced across pages) and a
real run_native drive proving the shutdown path actually cancels a
hanging task and invalidates it via refresh_timed_out.
Re-ran the review's own mutate_rev3.py (retargeted, unmodified
mutations): the two mutations flagged as regressions this round (G9, G10)
and G5 are now killed. G6 ("run_native never shuts down refresh tasks")
still survives -- verified redundant, not a real gap: run_native is only
ever invoked via asyncio.run() (both directly and through main()'s own
asyncio.run(execute())), and asyncio.run()'s own Runner.close() calls
_cancel_all_tasks() before shutting down the loop, cancelling every
outstanding task regardless of this module's own explicit cleanup. The
explicit _shutdown_corporate_action_tasks call remains valuable for the
exception-path case (finding, LOW #7, now fixed via the nested finally)
and for a hypothetical caller that reuses an already-running loop, not for
this survivor's literal scenario.
Recomputed source-hashes.json and the matching manifests/evidence.json
entries programmatically from the changed files' actual on-disk content;
registered the native-check record's own row updates.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 24, 2026
…ough NautilusTrader rc5 (audit gap #7) (#218) Replays paper trial 20260923g-main-passed (5 orders) through NautilusTrader 2.0.0rc5's simulated exchange on the same SIP quotes. Agreement is 3/5 at 0-50 ms latency and 5/5 from 70 ms, with a bisected flip at 69.217 ms: order 4 was marketable for 69.273 ms after submit, and order 5 depends on it. Engine settlement semantics, clock provenance and the cancel-pairing assumptions are disclosed. This is one trial and an agreement measurement, not a calibration. Reviewed across five rounds by Claude and Codex (#217). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins
deleted the
codex/late-native-tools-integration
branch
September 25, 2026 18:47
seathatflowsinourveins
added a commit
that referenced
this pull request
Sep 29, 2026
…er-bash parser (#516) * Add failing-first N1 probes and must-stay controls for M4 N1 is the #432 fixup3 residual: openers() does not track "$(" inside double quotes, and a quoted string that a shell runs is appended without its own heredoc resolution, so heredoc data reads as executed. Adds three tests through the existing measure/exports bridges on all three shell carriers plus fetchKind: - p1 (git commit -m "$(cat <<'EOF' ...)" with a line-start curl), its body variants with ', " and ( ), p2 (gh api in the body), p3 (bash -c 'cat <<EOF > x.sh ...') and a double-quoted p3 variant; - must-stay controls: an executed curl in "$( )", a shell heredoc in "$( )", an escaped substitution in a double-quoted run string, the command after a heredoc in "$( )" (guards against resolving one heredoc twice), the outer-shell heredoc in a double-quoted run string, unquoted $( ) and the "Subject (scope)" body; - the nested-quote case from design 1.5. At cf3fb72e the new tests fail 23 subtests (p1, the ' variant, p3 and the double-quoted p3: shell fetch 1 != 0 and fetchKind 'fetch'; p2: unclassifiable 1 != 0; nested quotes: 0 != 1 and fetchKind None); the must-stay test passes. Source: POSIX.1-2024 XCU 2.6.3 Command Substitution, checked with GNU bash 5.2.21 and dash. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add failing-first N1 tests for strings a shell runs The N1 fix analyzes a string that a shell runs (RUN_QUOTED) as that shell's input in both modes, instead of appending it verbatim unless inlineHttp is set. Two consequences get their own tests first: - in a double-quoted run string the outer shell removes the backslash before " (POSIX.1-2024 XCU 2.2.3), so `bash -c "echo \"a; curl ...\""` runs only echo: base reads a confirmed shell fetch; - a string the inner shell runs in turn is analyzed too, so the curl in `ssh host 'bash -c "curl ..."'` is confirmed: base counts it in neither the confirmed nor the unconfirmed fetches, so the gate's lower bound could not see it. Against the cf3fb72e kernel the test fails 7 subtests (1 != 0 and 0 != 1 on each carrier, fetchKind ['fetch', None]). GNU bash 5.2.21 and dash print `a; echo RAN` for the escaped form. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Resolve heredocs inside "$( )" and in strings a shell runs (N1) Fixes the #432 fixup3 residual N1 in child-usage.mjs. POSIX.1-2024 XCU 2.6.3: inside double quotes the text from "$(" to the matching ")" is tokenized recursively; 2.2.3: in a double-quoted word the outer shell removes the backslash before $ ` " \ and newline. GNU bash 5.2.21 and dash agree on every shape below. - step() is one frame machine (', ", `, and $( with depth) shared by phase 1 openers() and the phase 2 closeQuote()/matchParen(), so the phases agree on where a quote or substitution ends. With no frame open it reads exactly as the old openers(); only a double-quoted word gains $( and backquote frames, so a << inside "$( )" is found and its body resolved like any other heredoc. - executedTrace splits into resolveHeredocs (phase 1) and scanQuotes (phase 2). The naive end-of-quote scan becomes closeQuote; a "$( )" body in double-quoted data is spliced as "$(" + scanQuotes(body) + ")" with exact offsets (no second phase 1 over it). - A string a shell runs (RUN_QUOTED) is analyzed with the full executedTrace in both modes, a double-quoted one after unquoted() removes the outer shell's escapes at its own level. A shared set of resolved operator offsets keeps a heredoc the outer shell resolved in "$( )" from reading later lines as its body again. - NESTING_LIMIT (32) bounds recursion: deeper analyses read as data. Without it '"$(' x 3000 and a 3000-deep shell heredoc chain throw RangeError, which would abort a sweep. - The paren-free DOUBLE_QUOTED_DATA regex is removed; the scanners are linear character scanners (no new regex; CodeQL js/redos). The failing-first tests from 1d206a12 and 54299c44 now pass; the must-stay controls are unchanged. SHA256SUMS is regenerated in a later stage. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Document the N1 reading and add kernel unit cases The workflows README replaces the N1 limit paragraph with the rule the kernel now follows: "$( )" inside double quotes is read recursively (POSIX.1-2024 XCU 2.6.3), a string a shell runs is analyzed as that shell's input after the outer escapes are removed (XCU 2.2.3), and reading text as data moves a raw detector match to the unconfirmed count instead of dropping it. It states the remaining limits: backquoted spans inside double quotes, case patterns inside "$( )", the 32-level nesting bound, and text bash rejects as incomplete. tools/skill-usage/README.md notes that the legacy Python executed_text rule predates this reading and can disagree on these shapes. test-child-usage.mjs adds fetchKind and exact executedText cases for the N1 shapes and a nesting witness. Against the cf3fb72e kernel the three new cases fail (95/98); with NESTING_LIMIT removed the witness fails (97/98, RangeError); with the fix all 98 pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * State how the N1 reading moves the M4 lower bound The previous wording said reading text as data leaves the denominator of routed_share_lower_bound unchanged. That holds for raw detector matches only. Measured against the cf3fb72e kernel: - `echo "$(echo "a" && curl http://127.0.0.1:9/)"` moves from unconfirmed to confirmed loopback, which leaves the denominator; - `echo "$(\curl ...)"` gains a confirmed fetch no raw detector sees; - `echo "$(echo " ; \curl ...")"` loses a confirmation the old reading created (bash prints it as data), so the bound can rise. The README now states those three cases. It also adds a limit: a backquoted span outside quotes is still read as plain command text, so a quote it leaves open runs past its closing backquote. test-envelope (254/254), test-usage-receipts (18/18) and test-contract-mutations (74/74) pass; they read this README as routing_doc and contract_doc. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add failing-first CLI-lane, state and proxy-position fixtures Add the U1 design section 7 fixtures before the kernel change. The Python integration cases feed measureTranscript, aggregateMeasurements and the sweep CLI: the stack.json lane commands (:1003, :797, :579, :762, :519, :629, :280, :407), wrappers, compound commands and substitutions, runners, every mcporter call form, negatives (data, lookups, registrations, near names, gcm and serena-hooks), version and help, remote and unresolved programs, the call states, the rtk proxy population by command position with prefix_rule_calls, M3 and M4 by_carrier, aggregation with actors_with_success, and ID-free sweep output. The node suite adds commandInvocations unit cases. Sources: #381 AA-PLAN PR-A item 3 and M6/M14; POSIX.1-2024 XCU 2.9.1-2.9.4 and 2.6.3; rtk-ai/rtk@1d87b8e7 src/main.rs:68-90,708-713,3008-3042 and the installed rtk 0.50.0 (proxy -- and --skip-env bind, -v after proxy is the program); openclaw/mcporter@93e0916c src/cli.ts:115-364, src/cli/command-inference.ts:10-95, call-arguments.ts:79-233 and call-command.ts:114-173,309-348 (a first positional with = is the selector). Before the fix: 10 Python tests fail (102 subtests) and the 8 new node cases fail on the missing export. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Count CLI lanes and rtk proxy calls by command position Export commandInvocations: every simple command of the executed text (POSIX.1-2024 XCU 2.9.1-2.9.4, 2.6.3) is resolved past reserved words, assignments, wrappers (timeout, env, nice, stdbuf, command, exec, time, nohup, sudo 1.9.15p5, xargs), package runners and rtk proxy to a lane executable by exact basename. mcporter operations and a call's downstream server follow openclaw/mcporter@93e0916c (cli.ts, command-inference.ts, call-arguments.ts, call-command.ts); HTTP and stdio selectors read (http) and (stdio), and no host is emitted. rtk proxy follows rtk-ai/rtk@1d87b8e7 src/main.rs:68-90,708-713,3008-3042 and the installed rtk 0.50.0 (no shell runs the proxied program). One memoized reading per call drives carrierOf (M3 and M4 by_carrier), measurement.proxy (adds invocations, nested, in_ctx_code and prefix_rule_calls) and the new measurement.cli_lanes with per-call states (AA-PLAN "Successful" and M14; not_executed from observed transcript markers). aggregateMeasurements sums the fields and adds actors_with_success; the sweep limits gain the static limits. executedTrace takes an optional marks collector for ssh-run text and data-heredoc substitutions; its output is unchanged (0 differences in executedText and fetchKind over 24,489 inputs). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Read mcporter URL and stdio selector forms with char scanners Replace the character-class regexes in httpUrl and httpToolSelector, the path-prefix alternation for an ad-hoc stdio selector and the blank split of an inline npx server with character checks and the existing shellWords splitter. The brief asks for linear character scanners rather than regexes in this kernel (CodeQL js/redos flagged it before). Behaviour is unchanged: the rules still follow openclaw/mcporter@93e0916c src/cli/http-utils.ts:1-69, call-argument-values.ts:68-83 and ephemeral-target.ts:134-158, and the node suite passes unchanged (107/107). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Document CLI lanes, call states and the command-position proxy rule Add a CLI-lanes subsection to the PR-A measurement fields: the lane executables with the gcm and serena-hooks exclusions, the resolver (wrappers, runners, rtk proxy), each cli_lanes field, the mcporter operation and downstream-server precedence of mcporter v0.14.1, the call states with the observed not-executed markers and their M14 mapping, the static limits, and the changed meanings of proxy.calls, M3 by_carrier.rtk_proxy and m4.by_carrier with prefix_rule_calls kept for #369-era comparison. Record the decision that proxy.in_ctx_code stays outside M6 (Gate A plan M6 and rtk rows) and state that rtk_parts, proxy_parts and model_typed keep their prefix rules. The skill-usage README points Codex readers to the same fields and states that declined items and legacy exec headers still read succeeded until the adapter maps them. Sources: rtk-ai/rtk v0.50.0 src/main.rs:68-90,3008-3042; openclaw/mcporter v0.14.1 src/cli and docs/call-heuristic.md; mksglu/context-mode v1.0.169 src/exit-classify.ts:15-33; POSIX.1-2024 XCU 2.9. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add a failing case: exec -- reads as an unresolved program U1 design 2.2 reads only the POSIX exec form and marks any word that starts with '-' after exec as unresolved, because bash's exec options are not pinned. The shells disagree on `exec -- cmd`: GNU bash 5.2.21 runs cmd, while dash tries to execute "--" (probe on this host). The kernel currently takes `--` as the end of options and reads qmd. Observed before the fix: "exec -- qmd mcp" => ["qmd/qmd"]. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Read any dash word after exec as an unresolved program Follow U1 design 2.2: only the POSIX exec form is read, so a word that starts with '-' after exec, `--` included, leaves the program unresolved. bash 5.2.21 runs the command after `exec --` while dash executes "--" itself, so neither reading is safe to assume. The failing case from 1fcc9f23 now passes (node suite 107/107). Also cover the inline npx server of an mcporter call (openclaw/mcporter@93e0916c src/cli/ephemeral-target.ts:134-158): a --server value that is an npx command line is an ad-hoc stdio server, and another blank-holding name reads (other). These cases were added after the implementation, so they are coverage, not failing-first. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add failing-first Codex fixtures for shell call states U1 design 6 fixtures (i)-(vii) run rollout records through measure_codex_records, the Node bridge and the kernel's cli_lanes and proxy. Record shapes follow openai/codex rust-v0.157.1 (36650394): CommandExecutionItem and its snake_case status (protocol/src/items.rs :199-285), exec end states (core/src/tools/events.rs:529-573: a rejection is declined with exit -1) and the unified exec response header (core/src/tools/context.rs:524-575). A rollout output never carries the success flag (protocol/src/models.rs:2173-2182). Before the adapter change 6 of 8 subtests fail: - (iii) code-mode item declined: qmd succeeded 1, expected failed 1 and not_executed 1; the same for a direct call's declined item; - (iv) legacy exec_command with "Process exited with code 1": rtk_proxy succeeded 1, expected failed 1; - (v) "Process running with session ID 3": qmd succeeded 1, expected unknown 1; - (vi) no header and no item: qmd succeeded 1, expected unknown 1; - (vii) local_shell_call running stack.json:280: codebase-memory-mcp succeeded 1 and downstream codebase-memory succeeded 1, expected unknown 1 in both. (i) nested rtk proxy (proxy.calls 1, proxy.nested 1, succeeded 1) and (ii) a failed item already pass. Must-stay controls pass: an exit 0 header, an exit line inside the output, a completed item that beats a running header, a failed item with its header, and the aggregate passing cli_lanes and proxy through with actors_with_success. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Map declined items and exec headers to Codex call states U1 design 6: the Codex adapter now hands the kernel's callState what a rollout shows about a shell call's outcome, so cli_lanes stops reading declined or failed legacy commands as succeeded. - A declined CommandExecution (openai/codex rust-v0.157.1 36650394 core/src/tools/events.rs:562-573, exit -1) sets is_error and native_state "declined" on both result paths: the call's own function_call_output and the item's aggregated output. The kernel counts it failed and not_executed. - A Bash-mapped call (exec_command, shell_command, shell or local_shell_call, chosen by call kind, never by output text) with no persisted item state is read by the unified exec header of core/src/tools/context.rs:524-548, and only by the lines before "Output:", in a linear line scan with a strict ASCII exit code. Exit 0 is success, another exit is_error. A running line (it wins over an exit line, unified_exec/process_manager.rs:1066-1071), a header with neither line or text without the header sets native_state "unknown": a rollout output never carries success (protocol/src/models.rs:2173-2182), and legacy history mode keeps no item (rollout/src/policy.rs:94-112). No shell, shell_command or local-shell handler exists at that revision, so those outputs read unknown unless an item state decides; apply_patch's "Exit code: N" text is not read. - codex_call_name shares the MCP namespace rule between the call loop and the shell-call set, so an MCP tool named exec_command is not a shell call. The failing-first fixtures from 4a7aa75a now pass (the three Python suites: 160 tests OK). The header unit test (test_exec_header_state_reads_only_the_pinned_header) was written after the parser and is coverage, not failing-first. The lanes CLI still writes nothing to stderr, and the Codex aggregate passes cli_lanes and proxy through unchanged. tools/skill-usage/README.md replaces the stage B sentence that said these calls still read succeeded. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Regenerate workflow SHA256SUMS for the U1 kernel and tests The N1 fix and the CLI-lane reading changed child-usage.mjs and test-child-usage.mjs; stage C leaves the kernel byte-identical to 0c421c66 (sha256 acc7bb51...). Regenerated with the documented command, `sha256sum -- *.mjs *.js *.json > SHA256SUMS` in examples/claude-native/workflows, and `sha256sum --check --strict SHA256SUMS` passes for all 13 files. Only those two digests change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Re-register the U1 files in manifests/evidence.json The hot-file protocol (docs/lanes.md:94-103) puts the shared manifest edit in the branch's last commit and re-registers every changed file that the manifest already lists, with host_receipts.register_file. scripts/validate.py checks SHA-256 and byte counts for each of them and failed before this commit ("SHA-256 mismatch", "byte count mismatch" for the workflows README, child-usage.mjs, test-child-usage.mjs, tests/test_token_measurement.py and tools/skill-usage/README.md). The eight re-registered entries are the two READMEs, SHA256SUMS, child-usage.mjs, test-child-usage.mjs, skill_usage.py, test_skill_usage.py and test_token_measurement.py. Only their sha256 and bytes change, and validate.py now passes (status passed, 7986 hashed files). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add failing-first tests for the U1 pivot loader and M4 repairs Tests only, run against 7091a300 before any fix (U1 pivot brief D1, D6-D9, D11; GPT-6 findings #8, #9, #11, #12, #13; Claude review R1, R2, R3, R9). - D1 loader (test-child-usage.mjs, test_skill_usage.py): pin file contents, loadShellParser statuses (installed, not_installed, hash_mismatch), directory order, --shell-parser, commandInvocations null and cli_lanes parser_unavailable with proxy prefix_fallback, aggregate statuses, withShellTree freeing its tree, and the Node bridge awaiting the loader. - D6: shell_script quotes argv with shlex.join, and -c is a script only for a shell; exec_header_state bounds the exit code to 9 digits (GPT-6 #11, #12 inputs verbatim). - D7: outer-shell substitutions in a double-quoted run string are executed text (GPT-6 #8 verbatim), with single-quote, comment, backquote, sh, eval and ssh variants and must-stay controls that count once, not once per view. - D8: n = 8000..64000 doubling for unclosed ((, $((, "$( ((, and ( << E runs, at most 2.5x per doubling and under 150 ms at 64000. - D9: the stress check lets an exception fail it, includes an unquoted '$('.repeat(3000) input (NESTING_LIMIT), and a mutation control (a reader that throws for long input) shows the check now fails. - R1, R3: a heredoc after a closed "$( )" keeps its reader; an escaped blank, semicolon or newline before a # is not a comment start. Every expected value of the D7, R1 and R3 commands was run under GNU bash 5.2.21 and dash with a stub curl (and a stub ssh that runs sh on its stdin, the remote login shell): both shells agreed on every count. At 7091a300: node suite 115 passed, 31 failed (5 linear, 26 parser); D7 9 commands, R1 5, R3 3 fail with 0 != 1 on all three carriers; exec_header_state raises ValueError on 5000 digits; shell_script flattens argv. The old D9 assertion passes when 7 of 7 stress inputs throw. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Quote Codex argv elements and bound the exec header exit code U1 pivot brief D6 (GPT-6 findings #11 and #12). shell_script joined an argv array with spaces, so a metacharacter inside one element became shell syntax that was never executed: ["echo", "qmd; rtk proxy qmd status"] counted one qmd call and one rtk proxy call. It now joins with shlex.join (Python 3 shlex.join, the inverse of shlex.split), so each element is one quoted word. The [shell, -lc, script] form (openai/codex rust-v0.157.1 codex-rs/core/src/shell.rs) stays the script itself, but only when the program is a shell: -c of `echo -c ...` is one of echo's arguments. exec_header_state passed the exit code text to int() whenever it was ASCII digits, and CPython raises ValueError above 4,300 digits (int max str digits, docs.python.org/3/library/stdtypes.html#int-max-str-digits), which stopped the scan on one malformed header. The code is an i32 written in decimal (at most 10 digits with its sign, codex-rs/core/src/tools/context.rs:534-540); a code of more than 9 digits is no native header and reads (False, "unknown"). The D6 tests of the previous commit pass: exec_header_state on 5000 digits, the shell_script cases, and the three Codex bridge cases (local shell call, code-mode item, `echo -c`). The existing Codex state tests still pass. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Scan unclosed (( in linear time: one paren table per text (D8, R9) U1 pivot brief D8 (GPT-6 finding #9; Claude review R9). Every unclosed "((" or "$((" made step() look ahead to the end of its line (indexOf the newline, then count parentheses), so n = 8000, 16000 and 32000 of "((" took 246, 1218 and 4888 ms at 7091a300 (four times per doubling), and openers() filtered its whole cut list once per heredoc operator. parenClose(s) now matches every "(" of a text with its ")" in one stack pass per line (bash(1) ARITHMETIC EVALUATION: (( )) and $(( )) count every ( and ) up to the end of the line, which is classic parenthesis matching), and step() reads its answer from the table. The last eight texts stay memoized so the scans that alternate between a text and its slices do not rebuild it; executedTrace clears the memo when a top-level analysis ends. openers() finds each operator's bounds in the ascending cut list by binary search instead of filter and find. Output is unchanged: executedText (both modes) and fetchKind of the base kernel and this one agree on all 150,000 checks (the 243 covering commands, the 10,000 heredoc fuzz strings, and two seeded 20,000-string token corpora with and without heredocs; 0 differences, 0 throws). Timings at n = 8000, 16000, 32000, 64000: "((" 3/6/13/27 ms, "$((" 4/9/19/40 ms, unclosed (( inside "$( " 5/8/18/38 ms, "(<<E" runs 8/14/30/64 ms, each within 2.5x per doubling and under 150 ms at 64000. The D8 test allows 5 ms of timer noise on the 2.5x bound (a linear scan showed 7.6 -> 19.4 ms once) and takes the best of five. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Read a # after an escaped blank or metacharacter as data (R3) Claude review R3 of 3cb7c4f6. In a "$( )" frame step() treated a # as a comment start whenever the previous character was a word break, ignoring whether that break was escaped. `x="$(echo a\ #b)"; curl ...` therefore read `#b)"; curl ...` as a comment: closeQuote and matchParen ran to the end of the command and the `; curl` fell inside double-quoted data (shell_fetch 0, fetchKind null, one unconfirmed mention). bash(1) COMMENTS: a word beginning with # starts a comment, and an escaped blank, `;`, `(` ... or newline (a line continuation) does not end a word (QUOTING: a backslash escapes the next character). escapedAt(s, k) counts the backslashes before s[k]; an odd count means it is escaped. It runs only for a # that follows a word break, and reads one run of backslashes per such #, so the scan stays linear. Verified with GNU bash 5.2.21 and dash and a stub curl: for each of the six R3 inputs (escaped blank, escaped semicolon, backslash-newline, an even backslash count before a real comment, a real comment inside "$( )", and a top-level escaped blank) both shells ran the stub curl once, so each is one executed fetch. R3 test passes on all three shell carriers. Differential of the previous commit's kernel and this one over 90,243 inputs (243 covering commands, 10,000 heredoc fuzz strings, and two seeded 40,000-string token corpora, with and without heredocs; executedText in both modes and fetchKind): 228 differing inputs, every one containing an escaped word break before a # (0 unexplained). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Add failing-first R1 shapes and the --shell-parser refusal test Tests only. Extends the R1 cases of 62cb6935 with the shapes the fix has to cover beyond the reviewer's repro, each run under GNU bash 5.2.21 and dash with a stub curl (both shells agreed): - a quote inside the substitution: FOO="$(echo "a b")" bash <<'EOF' and the ssh variant, where the naive word split cuts the quoted word at the inner blank (curl ran once; with cat as the reader it ran 0 times); - a string that runs over lines before the operator: x="a\nb" bash <<'EOF' and the single-quoted variant, and a "$( )" that closes on a later line with text after its ")" (curl ran once each); - two closed "$( )" words, and controls that already read correctly at f83f8367 (a separator before the operator, a backquote span, a nested "$( )" with its own heredoc, cat as the reader). At f83f8367 the 11 executed shapes each fail on all three shell carriers (33 subtests: shell_fetch 0 != 1), the four controls pass. Also: an explicit --shell-parser that names a directory with no install exits 2 with the reason and no report (the --rtk-db and --exceptions convention), never a silent fallback to the default. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Bound each heredoc by the cuts of its own frame level (R1) Claude review R1 of 3cb7c4f6. openers() kept one cut list for every nesting level, so the ( and ) that a double-quoted "$( )" adds also bounded the heredoc operators outside it: in FOO="$(pwd)" bash <<'EOF' the operator's command started at the closing quote, its reader resolved to the empty string, and a body that bash reads as a script became data (shell_fetch 0, one unconfirmed mention, fetchKind null; real bash and dash run the stub curl once). step() now reports each cut with its frame (on.cut(i, frame): the "$(" frame that a ( opens and a ) closes, else the frame the unit sits in, undefined for the top level) and openers() keeps one ascending cut list per level; an operator is bounded by the cuts of the level on top of the stack, found by binary search. The cuts of a substitution bound the commands inside it and no longer the command that contains it. commandWords() reads a "$( )" or a backquoted span inside double quotes as one unit of its word (POSIX.1-2024 XCU 2.6.3: the substitution's text is tokenized on its own), so a quote or blank inside it does not split the word: FOO="$(echo "a b")" bash <<'EOF' names bash. R1 tests (62cb6935, 58502c60) pass on all three shell carriers, and the lane reading of FOO="$(pwd)" bash <<'EOF' with qmd in the body counts one qmd call. The shapes below are documented limits, pinned by their own test: a command that a quote or "$( )" carries over lines (x="$(cat <<'A' ... \n)" bash <<'B' and x="a\nb" bash <<'EOF') has its head on an earlier line, so this line alone gives the operator no reader and the body reads as data (fetch_mentions_unconfirmed 1, status incomplete: a possible fetch, never a lost one). A first attempt cut the line at the closing quote's word end; measured on the fuzz corpus it named the wrong owner as often as the right one (ssh host cat "a\nb" bash <<EOF), so it was dropped. Differential of f83f8367 and this commit over 90,243 inputs (243 covering commands, 10,000 heredoc fuzz strings, two seeded 40,000-string token corpora, without and with heredocs; executedText in both modes, fetchKind and the M4 confirmed and unconfirmed counts): 17 inputs differ, all with a quote, backquote or "$( )" before a heredoc operator (0 unexplained); no input lost mass (confirmed plus unconfirmed); 16 keep both counts, one moves a fetch from confirmed to unconfirmed. That one, bash -s -- "$(pwd)"bash <<EOF with an unquoted-delimiter body, now names bash as the reader (correct) and reads the body as source, where a substitution inside single quotes is data: the "$( )"-free twin reads C=0 U=2 at 7091a300 as well (real bash runs curl once: the outer shell expands an unquoted heredoc body before the reader sees it, a limit that predates R1 and is unchanged). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Read a double-quoted run string in two views: outer expansion (D7) U1 pivot D7, GPT-6 finding #8. N1 read a double-quoted string that a shell runs only as that shell's input, so a substitution the OUTER shell runs while it expands the word was lost whenever the inner shell reads it as data: a quoted heredoc, single quotes, a comment. bash -c "cat <<'EOF'\n$(curl https://example.org)\nEOF" went from fetchKind 'fetch' (cf3fb72e) to null with shell_fetch 0. bash 5.2.21 and dash run the stub curl once: the expanding shell runs each unescaped "$( )" and backquote before the shell it starts reads anything (POSIX.1-2024 XCU 2.2.3 Double-Quotes, 2.6.3 Command Substitution). Two views now: outer bodies first, each as commands of their own, then the inner shell's read of the string with each outer span replaced by a placeholder word (its output is unknown), so a substitution both views could see counts once. A backslash-escaped \$( ) or backquote is no outer span and stays the inner shell's (bash -c "echo \$(curl u)" runs curl once, in the inner shell). Single-quoted run strings have no outer expansion and are unchanged. outerSpans() is the one scanner of a double-quoted string's substitutions (escape pairs skipped, "$((" arithmetic not a span but the substitutions inside found, NESTING_LIMIT and unterminated spans as before); quotedData() now consumes it, so the data view and the run view cannot disagree. The ssh remote ranges (marks) cover the literal text between spans, not the local substitutions. D7 tests (62cb6935) pass on all three carriers: 13 executed, 5 inner and 2 twice-run commands, 2 data commands; the qmd lane test counts one call. Evidence, all against GNU bash 5.2.21 and dash with stub curl/wget/ssh (ssh runs its arguments with sh -c, or reads stdin as the remote login shell), counting stub calls against the kernel's confirmed count (shell_fetch + loopback): - compositional oracle, 360 commands (4 double-quoted run carriers x 10 contexts x 4 atoms, two-atom strings, single-quoted carriers, heredoc readers x delimiters x bodies): base 7091a300 301 exact, 59 under, 0 over; 43b9d947 329, 31, 0; this commit 353, 7, 0. The 7 left are one class, a shell-read heredoc with an unquoted delimiter whose body has a substitution inside single quotes (the outer shell expands the body first), next commit. - random valid-bash fuzz, 40,000 candidates, 7,883 accepted by bash -n and dash -n with equal stub-call counts under both (361 with a fetch): base, f83f8367, 43b9d947 and this commit agree on the same 7,789 (98.81%) and the same 94 disagreements (12 over: word glue and a trailing backslash; 82 under: remote ssh commands and glued words), all present at 7091a300. - differential of 43b9d947 and this commit over the covering commands and three fuzz corpora (90,243 inputs): 13 inputs lose confirmed mass, every one invalid bash in at least one shell; of the 24 differing inputs without HTTP-library or gh tokens, 21 are invalid, 2 valid inputs move closer to real bash and 1 (xcat &&bash -c "'$(curl u)'") counts a substitution that a failing && never reaches, which no static reading models. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Add failing-first tests: an unquoted heredoc body is expanded first Tests only. D7 for a heredoc body (U1 pivot D7, same defect class as GPT-6 #8): with an unquoted delimiter the shell that reads the command line expands the body before the shell that reads the heredoc sees it, so each unescaped substitution runs in the outer shell whatever its quotes or a comment say (bash(1) Here Documents; POSIX.1-2024 XCU 2.7.4), quotes are literal, and a backslash escapes only $ ` \ and a newline. 20 commands, each run under GNU bash 5.2.21 and dash with stub curl and ssh: curl ran once for 17 (a single-quoted substitution, a comment, nested quotes, a substitution over lines, a <<- body with tab indents, arithmetic, an ssh reader, a run string in the body, and controls that already read correctly), twice for one (an outer and an inner substitution), and never for two (a quoted delimiter passes the body as is). At 67a4acbb 9 of the 17 fail on all three shell carriers (shell_fetch 0 != 1, one unconfirmed mention: bash <<EOF with echo '$(curl u)', the FOO="$(pwd)" and ssh readers, a comment, nested quotes, a multi-line substitution, <<-, and arithmetic); the rest and the data commands pass. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Expand an unquoted heredoc body a shell reads before reading it The heredoc form of D7 (GPT-6 #8: the outer shell runs what it expands before the inner shell sees the text). With an unquoted delimiter the shell that reads the command line expands the body first: each unescaped $( ) and backquote runs there whatever its quotes or a comment say, quotes are literal in a body, and a backslash escapes only $ ` \ and a newline (bash(1) Here Documents; POSIX.1-2024 XCU 2.7.4). resolveHeredocs() analyzed a body that a shell (or in inlineHttp mode an interpreter) reads only as that reader's source, so bash <<EOF with echo '$(curl u)' read as data although real bash and dash run curl once, and so did a substitution in a comment, nested quotes inside it, one over lines, arithmetic, and a <<- body: the same 9 shapes as the failing-first commit b6f3b454. R1 made the FOO="$(pwd)" bash <<EOF form read this way too, so this also removes the accident that had hidden the limit there. The body now has two views, sharing D7's outerSpans/withoutSpans/outerBody: the outer spans first, as commands of their own (heredocs inside them resolved there), then the reader's source with each span replaced by the placeholder word and its own escapes applied (heredocText). A quoted delimiter is unchanged (nothing is expanded), and so is a data heredoc such as cat <<EOF, which keeps the substitutions its regular expression finds. The ssh remote ranges cover the body's literal text, not its local substitutions. The heredoc test of b6f3b454 passes on all three carriers (17 once, 1 twice, 2 data commands); tests.test_token_measurement passes (59). Evidence, real bash 5.2.21 and dash with stub curl/wget/ssh, kernel confirmed count against stub calls: - compositional oracle, 441 commands (D7's 360 plus escaped, comment and nested-quote bodies): base 7091a300 368 exact, 73 under; 67a4acbb 420, 21; this commit 441 of 441 exact, 0 under, 0 over. - grammar fuzz, 6,000 valid nested commands (statements, sequences, substitutions, backquotes, single and double quotes, run strings, heredocs with unique delimiters) that bash and dash ran with equal stub-call counts (3,049 with a fetch, 1,597 with a run string, 1,471 with a heredoc): base 5,788 exact (96.47%), 212 under, 0 over; 67a4acbb 5,970 (99.50%), 30 under, 0 over; this commit 5,975 (99.58%), 25 under, 0 over. All 25 are one older gap, a run string right after a backquote (echo `bash -c "curl u"`), whose text is invisible to the raw detector as well; fixed next. - random valid-bash fuzz, two seeds of 40,000 candidates (7,883 and 7,838 valid, 361 and 374 with a fetch): base, 67a4acbb and this commit agree on every input. - differential of 67a4acbb and this commit over 90,243 inputs: 924 differ, 20 with a lower confirmed count, of which 19 are invalid bash or valid only with a heredoc-at-end-of-file warning, and the one clean input (a backquote span that holds bash <<EOF) moves from 3 to 0 confirmed where real bash and dash run curl 0 times. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Add failing-first tests: a run string after an opening backquote Tests only. RUN_QUOTED anchors a shell, eval or ssh word at the text start, a blank, ; & | or (, but not at a backquote, so bash -c "curl u" inside `...` read as double-quoted data. The fetch was lost, and the raw detector, which anchors curl itself and not the quote, counted no possible fetch either (C=0 U=0). Found by the grammar fuzz of 2a33ef90: all 25 commands left under-counted there were this one shape. 8 commands run under GNU bash 5.2.21 and dash with stub curl and wget: the stub ran once for each (bash -c and sh -c with double or single quotes, eval, an ssh string, a backquote inside a single-quoted run string, an escaped backquote in a double-quoted one, and two controls that already read correctly: a blank after the backquote and $( ) in its place) and never for echo `echo "curl u"` (data). At 2a33ef90 the 6 backquote-adjacent shapes fail on all three shell carriers (shell_fetch 0 != 1, no unconfirmed mention: 19 failures with fetchKind). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Anchor a run string at an opening backquote too RUN_QUOTED and RUN_HTTP_CODE anchored a shell, eval, ssh or interpreter word at the text start, a blank, ; & | or (, but not at the backquote that opens a `...` substitution, so bash -c "curl u" inside backquotes read as double-quoted data: the fetch was lost and no possible fetch was counted either (C=0 U=0). bash(1) Command Substitution: a `...` body is shell text with its own command positions, as a $( ) body is; the anchor class now holds the backquote (one character in each regular expression; no new alternative, so no new backtracking). The 6 backquote shapes of the failing-first commit (66550644) pass on all three carriers; tests.test_token_measurement passes (60). Evidence, real bash 5.2.21 and dash with stub curl/wget/ssh, kernel confirmed count against stub calls: the 441-command oracle stays 441 exact; the grammar fuzz that found the gap is now exact on every input, 6,000 of 6,000 (3,049 with a fetch, 1,597 with a run string, 1,471 with a heredoc; 2a33ef90 5,975 exact, 7091a300 5,788, 96.47%) and, on a second seed, 11,999 of 11,999 (base 11,608, 96.74%, 391 under; 0 over throughout). Differential of 2a33ef90 and this commit over 90,243 inputs: 216 differ, 1 with a lower confirmed count, an unterminated string (invalid bash). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Add a pinned, verified tree-sitter-bash loader that fails closed (D1) U1 pivot D1. The CLI-lane reading is to run on tree-sitter-bash, the grammar OpenAI Codex uses for the same job (openai/codex rust-v0.157.1 36650394 codex-rs/shell-command/src/bash.rs; tree-sitter-bash = "0.25" at codex-rs/Cargo.toml:522, both read at the tag). This commit adds the seam and the gate; the AST reading itself is the next stage. shell-parser.pin.json pins the install: web-tree-sitter 0.27.0 and tree-sitter-bash 0.25.1 with their npm integrity values (checked against the registry's dist.integrity), upstream tags and commits (v0.27.0 6070dbfe = npm gitHead; v0.25.1 names a06c2e44 while the package was published from 80132668, five minutes earlier, recorded as a note), the sha256 of the six files the loader reads, and the install command (npm install --prefix <dir> --ignore-scripts --no-audit --no-fund --save-exact). loadShellParser(dir) resolves the directory as the argument (the new --shell-parser flag), then CHILD_USAGE_SHELL_PARSER, then the ecosystem tools directory under HOME, reads each pinned file once, compares its sha256 and the lockfile's version and integrity values with the pin, and executes only those bytes: the runtime module is imported from a private 0600 copy of the hashed bytes and the two wasm files are handed over as bytes, so the code that was hashed is the code that runs. Results carry no path: { ok, versions, wasm_sha256 } or { ok: false, reason } with not_installed (directory, a pinned file or the lockfile missing), hash_mismatch, load_error (an unreadable file, the pin file, init) and not_loaded (this process never awaited the loader, so a forgotten await is not mistaken for a missing install). A failed load leaves no parser, whatever an earlier call did; a successful one reuses the initialized runtime. withShellTree() frees every tree in a finally. Fail closed: commandInvocations() returns null without a parser and measurement reports cli_lanes { status: parser_unavailable, reason }, keeps the prefix rule for the rtk proxy carrier (proxy.rule prefix_fallback, invocations and in_ctx_code null) and never falls back to the text scanners for lane counting; M4 stays text based. A measured cli_lanes carries status measured and the parser record (versions and the two wasm sha256 values); aggregates report measured, incomplete (mixed actors, with the counts), parser_unavailable or not_measured, and sum proxy counters null-aware. The CLI awaits the loader; a --shell-parser that cannot be honored exits 2 with the reason (like --rtk-db), any other missing install only leaves cli_lanes unavailable. The Codex bridge awaits the loader too. Interim, until the AST reading lands: with a verified parser loaded, commandInvocations() still runs the scanner-based lane layer of 3cb7c4f6, so every existing lane fixture reads as before; the gate, the status fields and the seam (withShellTree) are final. Tests (62cb6935, 58502c60 first, run failing at 7091a300): the pin's contents; not_installed for an absent directory, a missing pinned file and a missing lockfile; hash_mismatch for one flipped byte in the bash wasm, a modified web-tree-sitter.js that is never imported, a modified package.json, another integrity value and another version in the lockfile; load_error for an unreadable file; directory order; --shell-parser; a good install loading again after failures; unavailable and mixed aggregates; the bridge in both states. The lane fixtures load the parser first and skip with a message on a host without an install. tests: node suite 149 passed; tests.test_token_measurement, test_skill_usage and test_child_usage_suite 175 passed. D9's stress check reports the TypeError the lane checks threw before the parser was loaded (threw: TypeError x8), which the old assertion would have passed. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Add failing-first timing tests for quote-heavy scripts (D8) Tests only. D8 asks the M4 text scanners for linear time; the unclosed "((" input of GPT-6 #9 is fixed (5e814331), but profiling a realistic 346 KB script (4,000 lines of echo "step N: $(date)" >> log; qmd search "term N" | head) at e5238f39 showed the same quadratic shape from another term: 51% of the time in the run-string regular expression and 33% in scanQuotes, both from one read of the whole text built so far per quote (the regex, and the string flattening that indexing a concatenation rope forces). A 42 KB script takes 47 ms, 84 KB 108 ms, 172 KB 467 ms, 346 KB 2,078 ms and 694 KB 8,236 ms a scan, and the kernel scans each shell call several times. Five doubling checks (n = 8000..64000, best of five, at most 2.5x per doubling plus 5 ms of noise, 64,000 under 1.5 s): a run of quoted words, a run of double-quoted substitutions in inlineHttp mode, a run of `bash -c "x"` strings, a run of words with a # inside, and a long script of echo, substitution and pipe lines through fetchKind. At e5238f39: 61, 230, 902, 5,861 ms; 168, 651, 3,728 ms; 991, 4,486 ms; 9, 32, 120, 459 ms; 89, 392, 1,930 ms (the series end where a run passes 1.5 s). All five fail. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Read only the tail of the built text at each quote (D8: linear scripts) scanQuotes() tested RUN_QUOTED (and in inlineHttp mode RUN_HTTP_CODE) against the whole text built so far at every quote. Reading out.s for that flattens the string concatenation and scans the full prefix each time, so a script with n quoted words cost n squared. CPU profile of a 346 KB script (4,000 lines of echo "step N: $(date)" >> log; qmd search "term N" | head) at e5238f39: 51% in the run-string regular expression, 33% in scanQuotes itself (the flattening), 10% in GC. The five timing tests of 9da57414 failed on it. scanQuotes now keeps the end of its output in `tail`, cut back to 512 characters whenever it passes 1,024, updated at every write (put and putTrace; quotedData builds into its own buffer first), and reads only that: the # test uses the last character of the tail, the emptiness test out.p.length. Once a cut has been made the two detectors run without their start-of-text alternative (a cut is no start of text). Limit, stated in the code: a run keyword more than 512 characters before its quote, after a text that has grown past 1,024, is missed and the string reads as data. Only a single ssh option word that long can put it there (ssh -oProxyCommand=<1500 characters> host "curl u": fetch before, null after; 100, 400, 520 and 700 characters read the same), since the words between a shell, eval or ssh and its string are short options and one host. Timings (best of five, ms), n = 8000, 16000, 32000, 64000: quoted words 61, 230, 902, 5,861 -> 8, 15, 31, 69; double-quoted substitutions (inlineHttp) 168, 651, 3,728 -> 15, 31, 67, 140; run strings 991, 4,486 -> 38, 83, 173, 358; words with # 9, 32, 120, 459 -> 3, 6, 14, 37; a long echo/substitution/pipe script through fetchKind 89, 392, 1,930 -> 18, 34, 71, 154. Realistic scripts, executedText: 84 KB 108 -> 37 ms, 346 KB 2,078 -> 100 ms, 694 KB 8,236 -> 205 ms, 1.4 MB 393 ms (2x per doubling); the node suite runs in 12 s instead of 41 s. Output is unchanged: executedText (both modes) and fetchKind of e5238f39 and this commit agree on all 1,084 shell texts of this repository (822 .sh files and fenced shell blocks of .md files, 1.66 MB, largest 82 KB, 83 over 4,000 characters), on the 243 covering commands and the three fuzz corpora (90,000 inputs), and on 600 synthetic long scripts (0 differences, 0 throws); the 441-command oracle stays 441 exact, and the 6,000-command grammar fuzz stays 6,000 exact (real bash 5.2.21 and dash with stub curl). Also fixed while rewriting the call: the ssh remote offset test listed the separators before RUN_QUOTED's keyword but not the backquote that 296c4f1d added, so in echo `ssh host "qmd get a"` the qmd call read as local; it now reads qmd/qmd remote (fixture added; the same input read local at e5238f39). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Flip one real byte in each pinned file in the loader tests The loader tests changed the bash grammar wasm by reading it as UTF-8 and appending a NUL, which rewrites the whole binary: the test passed because the file differed, not because one byte did. It now reads a Buffer and flips bit 0 of the middle byte in place (same length, so only the hash can tell), and does the same for each of the six pinned files in turn: every one gives hash_mismatch, leaves no parser (commandInvocations returns null, the scanners do not step in) and runs no unverified code (the modified web-tree-sitter.js sets a global if it is ever executed; it is not). The missing-file and missing-lockfile cases now also assert that no parser stays loaded. Node suite: 160 passed (149 before). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Add a failing-first test: a run keyword past the tail window Tests only. The brief's D8 says a work budget "marks the command unclassifiable for M4 instead of throwing". 1d77c34c bounds scanQuotes to the last 512 characters of its output, but an ssh option word long enough to push `ssh` out of the window (in a text past 1,024 characters) drops the run string silently: the fetch in ssh -oProxyCommand=<1500 characters> host "curl u" is read as data with no possible fetch counted either (shell_fetch 0, unclassifiable 0, unconfirmed 0), which a budget must not do. The test: option words of 100, 400 and 700 characters read as before (the fetch confirmed, unclassifiable 0); 1500 characters counts one unclassifiable operation (remote_fetches 1, no confirmed shell fetch); the count is once per command and only for the command that held the keyword (two later quoted commands add nothing); a long printf with hundreds of words and quotes but no run keyword adds none. At 1d77c34c the two 1500-character cases fail on all three shell carriers (unclassifiable 0 != 1: 6 failures); the 100, 400 and 700 cases and the printf control pass. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Count a run string the tail window lost track of as unclassifiable The brief's D8 fallback for input past a work budget is "marks the command unclassifiable for M4", not a silent miss. 1d77c34c bounds scanQuotes to the last 512 characters of its output; an ssh option word long enough to push `ssh` out of the window (in a text past 1,024 characters) dropped the run string with no trace: ssh -oProxyCommand=<1500 characters> host "curl u" read as shell_fetch 0, unclassifiable 0, unconfirmed 0. scanQuotes now notes, outside quotes, when a shell, eval, ssh or su word (and its blank) has been written since the last command separator (a last-letter pre-check keeps it off the hot path). When a cut has left the tail without that word and without a separator, the next quote that does not read as a run string is counted once in windowMisses, which countFetches adds to unclassifiable (an operation of unknown kind: remote_fetches 1, no confirmed shell fetch). The keyword is noted at write time because the discarded text cannot say whether a word was quoted: a first version that looked for the word in the discarded text also fired on 4 real scripts of this repository (adoption/bootstrap-linux.sh, bootstrap-macos.sh, launchd-agents.sh and a memory-stack README block), where separators inside quoted data are blanked and a word such as sh sat in a long quoted list. The test of 0b2b5778 passes on all three carriers (100, 400 and 700 characters read as before; 1500 counts one unclassifiable, once for the command, and adds nothing for later commands or for a long printf with no run keyword). The counter is zero on every corpus: complete m4 results of the pre-window kernel (e5238f39) and this one are identical on the 243 covering commands, 10,000 heredoc fuzz strings, two 40,000-string token corpora and all 1,084 shell texts of this repository (822 .sh files and fenced shell blocks, 1.66 MB), 2 tools each; the 600 synthetic long scripts agree; the 441-command oracle is 441 exact and the grammar fuzz 11,999 of 11,999 exact against real bash and dash. Realistic scripts, executedText: 84 KB 50 ms, 346 KB 125 ms, 694 KB 285 ms (2,078 and 8,236 ms at e5238f39). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Keep 150 ms for the brief input only; loosen the extra timing shapes The brief's D8 requires at most 2.5x per doubling and under 150 ms at n = 64,000 for '(('.repeat(n) + 'qmd'. The timing tests added alongside it (a run of $((, the same "$( ((, a run of ( << E, and the quote-heavy shapes) had inherited that absolute bound. Three of them run 64-80 ms at 64,000 against 150 ms on this host, thin headroom for a slower CI runner. Only the brief's input keeps 150 ms (it runs about 30 ms); every other shape keeps the 2.5x ratio (plus 5 ms of timer noise, best of five) and takes an absolute bound of 1.5 s at 64,000, which is still far below the quadratic behavior they replaced (the same shapes take 1.9 to 10 s and fail the ratio at 7091a300). Run against the base kernel (7091a300) with this test file, all ten timing shapes fail (for example (( 313, 1,219, 5,109 ms; quoted words 62, 239, 928, 6,211 ms), and against this head all ten pass (for example (( 3, 7, 17, 32 ms; quoted words 8, 15, 32, 71 ms; run strings 42, 95, 205, 416 ms). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Retry the doubling timing rounds so one pause cannot fail a linear scan A direct run of test-child-usage.mjs at da2d578d failed one assertion, "a run of unclosed (( at n = 8000, 16000, 32000, 64000 takes [4, 7, 27, 31] ms": the 32,000 reading was 27 ms where the scan takes about 15, so the ratio to the 7 ms before it (3.9x) broke the 2.5x (+5 ms) bound. The same suite had passed inside the Python run seconds earlier. A flaky timing check is a CI hazard, so the doubling measurement now takes the best of five per size and, when the ratio still fails, up to two more rounds and the elementwise minimum of all of them: a pause cannot fail a linear scan, while a quadratic one fails every round. Checks on the final file: six consecutive runs and two runs beside six busy CPU loops, all 160 passed; against the base kernel 7091a300 all ten timing shapes still fail (the retry does not hide quadratic behavior). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Add failing-first tests for the AST command-position layer (D10, D11) The GPT-6 review of the scanner-based lane reading (3cb7c4f6) found ten command-position defects (#1-#7, #10 here; #8, #9, #11-#13 were repaired earlier) and the Claude review five more (R4-R7). This commit holds their probes before any repair, so each one fails for its defect. - tests/test_token_measurement.py: one test per finding, with the GPT-6 inputs verbatim, asserting what the finding states (lane calls, proxy calls, exclusions, unresolved programs, downstream keys, call states), not a display string. Also parse_errors (D2), the closed name vocabulary (D5) and mcporter's help/version tokens (R6: checked against openclaw/mcporter@93e0916c src/cli.ts:137-145 and flag-utils.ts:34-41, the code already matches; the test pins it). - test-child-usage.mjs: a program name is emitted only for a lane executable or a name the reading interprets (GPT-6 #10). - tests/test_command_position_oracle.py (D10): a deterministic generator and a hand-written probe list run under REAL bash 5.2 with env -i, a PATH of logging stubs, a temporary HOME and cwd and a 5 s limit; the multiset of lanes that ran must equal the lanes commandInvocations reads. Against the scanner layer: 38 of 99 probes and 59 of 457 generated commands disagree with the run. Every expected value is the run of real bash and dash under stub executables (the oracle covers each input), not what this kernel says. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Read shell text with the tree-sitter-bash AST for CLI lanes (D2-D4) Replaces the scanner-based command-position layer (simpleCommands, resolveInvocation, analyzeScript) with a walk of the tree-sitter-bash tree, the grammar openai/codex rust-v0.157.1 codex-rs/shell-command src/bash.rs uses for the same job. The option tables, runner maps and mcporter grammar are unchanged and now operate on word values. - D2: every `command` node counts wherever it sits (substitutions inside strings, arithmetic, arrays and unquoted heredoc bodies, subshells, compound and control statements, function bodies); array elements, case patterns, [[ ]] operands, function names and the words of any non-lane program never do. ERROR nodes and what is under them are skipped and the call counts once in cli_lanes.parse_errors. - D3: wordValue() returns a literal or null: unquoted escapes, '...', the double-quote rule (a backslash goes only before $ ` " \ and newline), $'...' decoded as bash does, concatenations; any expansion, glob, brace expansion is unknown. A tilde prefix keeps the word known for program identity (only the directory changes: ~30 lane invocations by ~/path in 136,361 real commands) but not as a script. - D4: bash|sh|dash|zsh|ksh -c reads the first operand after the options (-n, -D and -o noexec run nothing; `--`, `-` end the options); a shell with no -c reads standard input, so a heredoc or here-string attached to the last command of a list, pipeline or negation, through env, timeout, nice, nohup, stdbuf, rtk proxy (not xargs, which consumes it), has its body read; eval and ssh join their words as the callee does; rtk proxy resolves its program with exec semantics (command and exec run nothing, time is unresolved, an expansion in its one argument is unresolved). - Backquotes and outer substitutions of an unquoted heredoc body and of a double-quoted script are read from the raw text with the kept M4 helpers (tree-sitter-bash leaves backquotes in heredoc_content and mis-reads some bodies); a script with an expansion is read with a placeholder. - The grammar folds the next line into a command in some shapes (an unquoted == before ;, &&, | or a newline; a trailing blank before a newline; `<<'EOF' 2>&1 | tail`): 732 of 136,361 real commands with no error and 163 with one. A separator ERROR or an unescaped newline inside one command's words starts a new command. - simple counts assignment-only statements, declarations and each loop, if and case as well as commands (POSIX.1-2024 XCU 2.9.1, 2.9.4), so the ambiguity of existing fixtures is unchanged. Fixtures: EMPTY_CLI gains parse_errors (D2), and `ssh host bash -s <<EOF` now also reads the remote bash as a command of its own (ssh(1): the words after the destination run remotely). Every other existing lane fixture passes unchanged. Against real bash 5.2: the oracle agrees on 457 of 457 generated commands (581 lane runs) and 99 of 99 probes. Over 136,361 distinct real commands the old and new readings agree on 136,167; the differences are old false positives, unknown wrappers (setsid, bwrap) and grammar defects. Still failing here by design, fixed next: D5 vocabulary (5 tests) and R7 states (7 tests); the node suite fails only on the D5 program-name check. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Re-expect the closed name vocabulary in the lane fixtures (D5) U1 pivot D5 and GPT-6 #10: commandInvocations returns `program` only for a lane executable or a name this reading interprets itself (a wrapper, a shell, eval, ssh, rtk), and an mcporter downstream key is emitted only for the stack's own servers; every other name-shaped string (a host, an id, a package, `linear`) reads null or (other). The old expectations that showed `-/echo`, `-/git`, `-/pytest` or `@linear` are wrong under D5, so they are changed here, before the code: 7 node checks and 3 Python tests fail against the walker of 1d35df12 with the observed values recorded in the log (`program` still holds "my-private-host.example", the downstream key still holds "linear" and "call_PRIVATE"). The selector-form coverage the `linear` cases gave is kept with servers of the stack (`socraticode`, `serena`, `jcodemunch`, `ai-memory`, `qmd`, `headroom`, `context-mode`, `codebase-memory`), read by name in every form. Sources of the closed set: manifests/stack.json:280 (codebase-memory), :407 (context-mode), :629 and :966 (socraticode); the rest is the coordinator's list in the pivot brief. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Emit only fixed names: a closed program set and mcporter server set (D5) GPT-6 #10: `program` carried any name-shaped basename (a host, an id, a package) and an mcporter downstream key any name that passed a character class, so the no-id, no-host output contract held only by luck. - `program` is set only for a lane executable or a name the reading interprets itself (the wrappers, the shells, eval, ssh, rtk); every other program reads null. - An mcporter downstream key is one of codebase-memory, context-mode, jcodemunch, serena, socraticode, qmd, headroom or ai-memory (the stack's servers: manifests/stack.json:280, :407, :629, :966 and the coordinator's list), else (other), (http), (stdio) or (unresolved). The tests of the previous commit now pass (node 162 of 162; the Python name-vocabulary test and the three re-expected fixtures). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Expect an interrupted counter in lane and downstream rows (R7) Claude review R7: a Claude Code result whose toolUseResult.interrupted is true read as succeeded. The command started and was cut short, so it is a state of its own. The shared fixture helpers (lane_row, downstream_row and the Codex cli_lane_row and downstream row) now carry `interrupted: 0`; the lane and downstream rows of the kernel do not have the key yet, so the tests that compare whole rows fail (see the log of this run) until the next commit adds the counter and the state. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Read not-executed calls where the client records them; add interrupted (R7) Claude review R7: the marker "The user doesn't want to proceed" was matched against the row's toolUseResult, where it was never observed; it opens the result CONTENT. Observed on this host's Claude Code transcripts (122,648 Bash results in 4,355 files, count-only): the rejection content (25 rows, toolUseResult "User rejected tool use"), a row-level toolDenialKind on every client denial (permission-rule 750, user-rejected 46, cancelled 1, automode-unavailable 1), the classifier text "The server-side auto mode classifier gave no verdict" (its own text says the check failed, not the action), and 209 rows of a host hook's refusal text ("This agent is isolated") that carry no toolDenialKind and stay failed (open: its meaning is the hook's, not the client's). - callState: not_executed when the row has a toolDenialKind, or the content opens with <tool_use_error>, "The user doesn't want to proceed" or the classifier text, or toolUseResult names a hook denial, a permission denial or a rejection. - toolUseResult.interrupted === true is its own state, `interrupted` (none of 25,506 observed object results, so it is a synthetic fixture), with a counter in lane and downstream rows; it no longer reads succeeded. The R7 test of the first commit and the fixtures of the previous commit pass (measurement and Codex modules: 188 tests). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Remove the lane-only marks plumbing from the M4 text machinery The scanner-based lane reading asked executedTrace for `marks` (the raw range of each ssh body and the offset of each command substitution of a data heredoc body). The AST walker reads both from the tree and the raw text, so nothing passes marks any more: the parameter and every branch that wrote to it (resolveHeredocs, scanQuotes, outerBody, quotedData) and rawRange are removed. The M4 readings are unchanged: the complete m4 object of measureTranscript (Bash and ctx shell carriers), executedText in both modes and fetchKind are identical between the kernel before the walker (f1ed98ac) and this one over 136,361 distinct real commands, the 243 covering commands, and 2 x 40,000 fuzz inputs (0 differences, 0 exceptions). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Test the recovery of commands the grammar joined to the previous one The walker of 1d35df12 starts a new command at a separator ERROR or an unescaped newline inside one command's words, because tree-sitter-bash 0.25.1 reads what follows a command as its arguments in these shapes (measured on this host's transcripts: 732 of 136,361 distinct real commands with no error and 163 with one): an unquoted == or =~ before ;, &&, | or a newline, a redirection after a here-document operator followed by | or ;, and a multi-line shape found in the real corpus. This commit adds the tests that pin it: - tests/test_command_position_oracle.py: 12 recovery probes whose expected lanes are what REAL bash 5.2 ran (12 of 12 agree; 10 carry a parse error, 2 none). - tests/test_token_measurement.py: the same shapes through measureTranscript, with parse_errors counted once for the calls whose tree has an ERROR, and two controls that are not boundaries (a backslash-newline continuation and a comment). Mutation control (recovery disabled in a scratch copy of the kernel): all 12 oracle probes fail (real bash ran the lane, the kernel read none or one fewer, e.g. `echo ==; qmd get a` ran qmd, read []) and 9 of the 12 measurement cases fail with the observed values ({} != {'qmd': 1}). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Add failing-first tests: a line that begins with a backslash Found by the 14,461-command run of the oracle generator over more seeds against real bash: a line that begins with a backslash (`\ls`, the alias bypass; `\markitdown a.json`) reaches tree-sitter-bash 0.25.1 as a word that begins with the newline, so the line joins the command before it, and when it is the first line of a here-document body the grammar leaves the body node without it and adds its words to the operator's arguments (which hid that the owner is a shell). A here-string that precedes such a line was fed to the last command of the joined node instead of its own. Five shapes are added to the oracle's recovery probes and to the measurement recovery test; real bash ran every lane in them. Against the kernel of 907ac35f the two tests fail with the observed values (recorded in the log of this run), e.g. `{ timeout 5 repomix` + newline + `\markitdown --flag; }` read {'repomix': 1} where bash ran both, and the shell heredoc whose first body line starts with a backslash read {}. Also hardens the oracle generator: it drops a text bash rejects (`bash -n`), keeps a function definition and its call in one group, does not put `!` or a heredoc on the right of a pipe (`x | ! y` is not POSIX; a reader that ignores its pipe kills the writer with SIGPIPE), and writes data heredocs to /dev/null. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> * Read a line that begins with a backslash as a command of its own Repairs the shapes of the previous commit's tests: - commandSegments starts a new command at a word that begins on a new line (the grammar puts the newline of a `\cmd` line at the start of its word) and drops that newline from the word's value. - A here-document operator's arguments are only those on its own l…
5 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Integrate the reviewed upstream-native-tool additions while retaining the published GPU trial and native recovery results. The three imported receipts now explicitly apply only to their source Linux/WSL host; they do not imply Mac or VelaNext installation. Resolve catalog/explorer conflicts structurally and pin generated source links to the ancestry-preserving merge.
The catalog now contains 57 component records, 61 receipts, 590 integrity records, 505 repository identities and 1,055 typed references. Native recipes cover Beads, the Agent Skills reference validator and otel-tui, with separate ai-memory maintenance and Linux zizmor evidence. Existing no-capture policy, optional GPU status and savings limitations remain.
Validation: 106 focused tests passed with a resolved macOS temporary path; full native metadata suite 429 tests passed with 47 environment-dependent skips; source/catalog/explorer and secret checks passed. An initial cross-project Python environment incorrectly activated optional DuckDB tests without pytz; the failure is retained and the suite was rerun using the documented standalone interpreter. No model calls or host upgrades were performed by this integration.