Skip to content

Integrate native tool evidence with explicit source-host scope - #7

Merged
seathatflowsinourveins merged 5 commits into
mainfrom
codex/late-native-tools-integration
Sep 20, 2026
Merged

seathatflowsinourveins merged 5 commits into
mainfrom
codex/late-native-tools-integration

Conversation

@seathatflowsinourveins

Copy link
Copy Markdown
Owner

Integrate the reviewed upstream-native-tool additions while retaining the published GPU trial and native recovery results. The three imported receipts now explicitly apply only to their source Linux/WSL host; they do not imply Mac or VelaNext installation. Resolve catalog/explorer conflicts structurally and pin generated source links to the ancestry-preserving merge.

The catalog now contains 57 component records, 61 receipts, 590 integrity records, 505 repository identities and 1,055 typed references. Native recipes cover Beads, the Agent Skills reference validator and otel-tui, with separate ai-memory maintenance and Linux zizmor evidence. Existing no-capture policy, optional GPU status and savings limitations remain.

Validation: 106 focused tests passed with a resolved macOS temporary path; full native metadata suite 429 tests passed with 47 environment-dependent skips; source/catalog/explorer and secret checks passed. An initial cross-project Python environment incorrectly activated optional DuckDB tests without pytz; the failure is retained and the suite was rerun using the documented standalone interpreter. No model calls or host upgrades were performed by this integration.

@seathatflowsinourveins
seathatflowsinourveins merged commit d4d3084 into main Sep 20, 2026
4 checks passed
seathatflowsinourveins added a commit that referenced this pull request Sep 23, 2026
…okenizer, exact hostname) (#104)

* Fix CodeQL first-analysis alerts: href scheme allowlist, case-insensitive tag scan, exact hostname check

Resolves the 5 real defects from the repository's first CodeQL default-setup
analysis (commit 168a3a8, alerts 1/3/4/5/8), re-located at base 796f759 since
PR #96 changed the generated explorer between the two:

- js/xss-through-dom (docs/ecosystem/template.html:145): link() now builds
  href through a safeHref() helper that only allows http:/https: URLs,
  blocking a javascript:-URI href from catalog data.
- py/bad-tag-filter (scripts/build_ecosystem.py:704,
  tests/test_ecosystem_manifest.py:231, tests/test_claude_repository_evidence.py:147):
  add re.IGNORECASE so an injected uppercase <SCRIPT> tag is still counted
  into the page's inline-script CSP hash instead of silently bypassing the
  single-script precondition.
- py/incomplete-url-substring-sanitization (tests/test_lifecycle_capture.py:103):
  replace the "sec.gov" in url substring check with an exact
  urlparse(url).hostname comparison.

The remaining 5 alerts (2, 6, 7, 9, 10) are false positives / test-only
synthetic-secret fixtures; their justification is recorded for the
coordinator to apply via the GitHub API in
docs/decisions/2026-09-22-codeql-first-analysis.md and the sibling
codeql-dismissals.json, not dismissed here.

manifests/evidence.json's sha256/bytes entries for the five touched files
are refreshed so scripts/validate.py (run by validate.yml) still passes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Correct alert 7/9 location prose and add an executed safeHref check (review fixes)

Resolves the independent reviewer's three medium findings on 65f73e2:
alerts #7 and #9's location descriptions and dismissal comments pointed at
the wrong code after the 796f759 re-location (both now cite the actual
CodeQL sink); alert #1's "node -e smoke check" claim named no runnable
command, so it is replaced with an executed unittest that runs the
committed safeHref helper (extracted verbatim from template.html) under
Node and asserts javascript:/data:/vbscript:/file:/mailto: are rejected
while http(s) and relative URLs pass through, and alert #1 is relabeled
defense-in-depth given build_ecosystem.py's public_url() is the primary
control. Refreshes manifests/evidence.json for the two touched files so
validate.py stays green. Re-running the full suite three times confirms
the reviewer's flagged skip-count (338) is deterministic, not a flake.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Replace the <script> regexes with stdlib HTML tokenizers (CodeQL py/bad-tag-filter)

re.I alone would likely leave py/bad-tag-filter open (it also flags missed
end-tag variants such as </script >). The build script and both tests now use
html.parser, which tokenizes like a browser; on the real generated pages the
parsers return the identical single body the regexes returned. Corrects the
record's alert #1 control description (loopback_url also admits http loopback
links) and the skip-count claim.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Fail loudly where html.parser and a browser disagree on <script>

Review follow-up: a self-closing <script/> or a script after <!--> (Python 3.12)
was invisible to InlineScripts. render_from_data now requires no self-closing
script and requires the parser's script-start count to equal the raw
"<script" count; the CSP test asserts the same count. The record's loopback_url
and browser-equivalence wording is corrected.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: seathatflowsinourveins <234074349+seathatflowsinourveins@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins pushed a commit that referenced this pull request Sep 24, 2026
… found

Both independent reviews confirmed round 3 mostly holds (all 24 earlier
mutations still killed) but independently found the same new regression:

- HIGH: ambiguity erased a confirmed action. AlpacaCorporateActionsSource.
  fetch() used to overwrite a symbol's whole record list with
  LOOKUP_AMBIGUOUS whenever ANY row for that symbol lacked a governing
  date, discarding valid records collected in the SAME response (e.g. a
  confirmed split plus an undated dividend for the same symbol).
  CorporateActionMonitor._apply_success also overwrote a previously
  confirmed record with a later per-symbol LOOKUP_AMBIGUOUS/LOOKUP_FAILED
  result. Fixed at both layers: fetch() only ever falls back to the
  sentinel for a symbol with ZERO valid records of its own in that
  response; _apply_success degrades (keeping the last good record)
  instead of overwriting whenever the new per-symbol result is a
  sentinel.
- MEDIUM: raw rows with no identifiable symbol used to silently count as
  clear. A row (mapped or unmapped bucket) naming no symbol at all now
  marks every requested symbol ambiguous.
- MEDIUM: uncertainty discovered after the final rebalance() tick could
  still end held_overnight. The outcome's corporate_action_guard summary
  is now recomputed fresh (_final_corporate_action_guard_summary) against
  the guard's current evaluate() and the truly final held positions,
  immediately before the outcome is built, instead of trusting the
  strategy's last-tick snapshot.
- LOW: shutdown errors could bypass refresh-task cleanup -- session.stop()/
  port.stop()/the task wait are now in their own nested try, with
  _shutdown_corporate_action_tasks in its own guaranteed finally.
- LOW: the private _get_marketdata call is replaced with the public
  client.get() and this wrapper's own explicit next_page_token pagination
  (matching blueprints/us-equities/alpaca-historical's own bridge pattern
  for this endpoint), with region=us and a page-count ceiling that fails
  closed if a continuation token survives it.
- LOW: one unmapped-type row used to fail the fetch for every requested
  symbol. An unmapped bucket now degrades only the symbol(s) its own rows
  actually name (Alpaca API reference, corporateactions-1, fetched
  2026-09-24: partial_call/reorganization/capital_gains_distribution
  documented as additional types; data_quality accepts complete/all).
- LOW (measured): a hung refresh thread used to delay run_native's own
  result-save/account-lock release by the hang's full duration, because
  asyncio.run()'s cleanup joins every asyncio.to_thread/default-executor
  thread. The worker now runs on a daemon thread bridged back via
  loop.call_soon_threadsafe; asyncio.run() never waits for it (measured:
  a 4s hang delayed exit by ~4s before, ~0.1s after).
- NIT: a timeout processed immediately after a successful commit for the
  same generation is now ignored (a new _committed_generation check),
  instead of degrading an already-fresh result for the retry interval.
- Wording: corrected the "never overlapped"/"bounded fetch" claims to
  state exactly what actually bounds each thing, and the native-check
  record's "2026 is the only calendar year" claim (2027 is also covered).

Adds targeted tests for every finding above, including tests that go
through the real HTTP-mocked SDK pagination path (a two-page response
with a real next_page_token, and the cap enforced across pages) and a
real run_native drive proving the shutdown path actually cancels a
hanging task and invalidates it via refresh_timed_out.

Re-ran the review's own mutate_rev3.py (retargeted, unmodified
mutations): the two mutations flagged as regressions this round (G9, G10)
and G5 are now killed. G6 ("run_native never shuts down refresh tasks")
still survives -- verified redundant, not a real gap: run_native is only
ever invoked via asyncio.run() (both directly and through main()'s own
asyncio.run(execute())), and asyncio.run()'s own Runner.close() calls
_cancel_all_tasks() before shutting down the loop, cancelling every
outstanding task regardless of this module's own explicit cleanup. The
explicit _shutdown_corporate_action_tasks call remains valuable for the
exception-path case (finding, LOW #7, now fixed via the nested finally)
and for a hypothetical caller that reuses an already-running loop, not for
this survivor's literal scenario.

Recomputed source-hashes.json and the matching manifests/evidence.json
entries programmatically from the changed files' actual on-disk content;
registered the native-check record's own row updates.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
seathatflowsinourveins added a commit that referenced this pull request Sep 24, 2026
…ough NautilusTrader rc5 (audit gap #7) (#218)

Replays paper trial 20260923g-main-passed (5 orders) through NautilusTrader 2.0.0rc5's simulated exchange on the same SIP quotes. Agreement is 3/5 at 0-50 ms latency and 5/5 from 70 ms, with a bisected flip at 69.217 ms: order 4 was marketable for 69.273 ms after submit, and order 5 depends on it. Engine settlement semantics, clock provenance and the cancel-pairing assumptions are disclosed. This is one trial and an agreement measurement, not a calibration. Reviewed across five rounds by Claude and Codex (#217).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@seathatflowsinourveins
seathatflowsinourveins deleted the codex/late-native-tools-integration branch September 25, 2026 18:47
seathatflowsinourveins added a commit that referenced this pull request Sep 29, 2026
…er-bash parser (#516)

* Add failing-first N1 probes and must-stay controls for M4

N1 is the #432 fixup3 residual: openers() does not track "$(" inside
double quotes, and a quoted string that a shell runs is appended
without its own heredoc resolution, so heredoc data reads as executed.

Adds three tests through the existing measure/exports bridges on all
three shell carriers plus fetchKind:
- p1 (git commit -m "$(cat <<'EOF' ...)" with a line-start curl), its
  body variants with ', " and ( ), p2 (gh api in the body), p3
  (bash -c 'cat <<EOF > x.sh ...') and a double-quoted p3 variant;
- must-stay controls: an executed curl in "$( )", a shell heredoc in
  "$( )", an escaped substitution in a double-quoted run string, the
  command after a heredoc in "$( )" (guards against resolving one
  heredoc twice), the outer-shell heredoc in a double-quoted run
  string, unquoted $( ) and the "Subject (scope)" body;
- the nested-quote case from design 1.5.

At cf3fb72e the new tests fail 23 subtests (p1, the ' variant, p3 and
the double-quoted p3: shell fetch 1 != 0 and fetchKind 'fetch'; p2:
unclassifiable 1 != 0; nested quotes: 0 != 1 and fetchKind None); the
must-stay test passes. Source: POSIX.1-2024 XCU 2.6.3 Command
Substitution, checked with GNU bash 5.2.21 and dash.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add failing-first N1 tests for strings a shell runs

The N1 fix analyzes a string that a shell runs (RUN_QUOTED) as that
shell's input in both modes, instead of appending it verbatim unless
inlineHttp is set. Two consequences get their own tests first:
- in a double-quoted run string the outer shell removes the backslash
  before " (POSIX.1-2024 XCU 2.2.3), so `bash -c "echo \"a; curl
  ...\""` runs only echo: base reads a confirmed shell fetch;
- a string the inner shell runs in turn is analyzed too, so the curl
  in `ssh host 'bash -c "curl ..."'` is confirmed: base counts it in
  neither the confirmed nor the unconfirmed fetches, so the gate's
  lower bound could not see it.

Against the cf3fb72e kernel the test fails 7 subtests (1 != 0 and
0 != 1 on each carrier, fetchKind ['fetch', None]). GNU bash 5.2.21
and dash print `a; echo RAN` for the escaped form.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Resolve heredocs inside "$( )" and in strings a shell runs (N1)

Fixes the #432 fixup3 residual N1 in child-usage.mjs. POSIX.1-2024 XCU
2.6.3: inside double quotes the text from "$(" to the matching ")" is
tokenized recursively; 2.2.3: in a double-quoted word the outer shell
removes the backslash before $ ` " \ and newline. GNU bash 5.2.21 and
dash agree on every shape below.

- step() is one frame machine (', ", `, and $( with depth) shared by
  phase 1 openers() and the phase 2 closeQuote()/matchParen(), so the
  phases agree on where a quote or substitution ends. With no frame
  open it reads exactly as the old openers(); only a double-quoted
  word gains $( and backquote frames, so a << inside "$( )" is found
  and its body resolved like any other heredoc.
- executedTrace splits into resolveHeredocs (phase 1) and scanQuotes
  (phase 2). The naive end-of-quote scan becomes closeQuote; a "$( )"
  body in double-quoted data is spliced as "$(" + scanQuotes(body) +
  ")" with exact offsets (no second phase 1 over it).
- A string a shell runs (RUN_QUOTED) is analyzed with the full
  executedTrace in both modes, a double-quoted one after unquoted()
  removes the outer shell's escapes at its own level. A shared set of
  resolved operator offsets keeps a heredoc the outer shell resolved
  in "$( )" from reading later lines as its body again.
- NESTING_LIMIT (32) bounds recursion: deeper analyses read as data.
  Without it '"$(' x 3000 and a 3000-deep shell heredoc chain throw
  RangeError, which would abort a sweep.
- The paren-free DOUBLE_QUOTED_DATA regex is removed; the scanners are
  linear character scanners (no new regex; CodeQL js/redos).

The failing-first tests from 1d206a12 and 54299c44 now pass; the
must-stay controls are unchanged. SHA256SUMS is regenerated in a
later stage.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Document the N1 reading and add kernel unit cases

The workflows README replaces the N1 limit paragraph with the rule the
kernel now follows: "$( )" inside double quotes is read recursively
(POSIX.1-2024 XCU 2.6.3), a string a shell runs is analyzed as that
shell's input after the outer escapes are removed (XCU 2.2.3), and
reading text as data moves a raw detector match to the unconfirmed
count instead of dropping it. It states the remaining limits:
backquoted spans inside double quotes, case patterns inside "$( )",
the 32-level nesting bound, and text bash rejects as incomplete.

tools/skill-usage/README.md notes that the legacy Python executed_text
rule predates this reading and can disagree on these shapes.

test-child-usage.mjs adds fetchKind and exact executedText cases for
the N1 shapes and a nesting witness. Against the cf3fb72e kernel the
three new cases fail (95/98); with NESTING_LIMIT removed the witness
fails (97/98, RangeError); with the fix all 98 pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* State how the N1 reading moves the M4 lower bound

The previous wording said reading text as data leaves the denominator
of routed_share_lower_bound unchanged. That holds for raw detector
matches only. Measured against the cf3fb72e kernel:
- `echo "$(echo "a" && curl http://127.0.0.1:9/)"` moves from
  unconfirmed to confirmed loopback, which leaves the denominator;
- `echo "$(\curl ...)"` gains a confirmed fetch no raw detector sees;
- `echo "$(echo " ; \curl ...")"` loses a confirmation the old
  reading created (bash prints it as data), so the bound can rise.
The README now states those three cases. It also adds a limit: a
backquoted span outside quotes is still read as plain command text,
so a quote it leaves open runs past its closing backquote.

test-envelope (254/254), test-usage-receipts (18/18) and
test-contract-mutations (74/74) pass; they read this README as
routing_doc and contract_doc.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add failing-first CLI-lane, state and proxy-position fixtures

Add the U1 design section 7 fixtures before the kernel change. The
Python integration cases feed measureTranscript, aggregateMeasurements
and the sweep CLI: the stack.json lane commands (:1003, :797, :579, :762,
:519, :629, :280, :407), wrappers, compound commands and substitutions,
runners, every mcporter call form, negatives (data, lookups,
registrations, near names, gcm and serena-hooks), version and help,
remote and unresolved programs, the call states, the rtk proxy
population by command position with prefix_rule_calls, M3 and M4
by_carrier, aggregation with actors_with_success, and ID-free sweep
output. The node suite adds commandInvocations unit cases.

Sources: #381 AA-PLAN PR-A item 3 and M6/M14; POSIX.1-2024 XCU 2.9.1-2.9.4
and 2.6.3; rtk-ai/rtk@1d87b8e7 src/main.rs:68-90,708-713,3008-3042 and
the installed rtk 0.50.0 (proxy -- and --skip-env bind, -v after proxy
is the program); openclaw/mcporter@93e0916c src/cli.ts:115-364,
src/cli/command-inference.ts:10-95, call-arguments.ts:79-233 and
call-command.ts:114-173,309-348 (a first positional with = is the
selector). Before the fix: 10 Python tests fail (102 subtests) and the
8 new node cases fail on the missing export.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Count CLI lanes and rtk proxy calls by command position

Export commandInvocations: every simple command of the executed text
(POSIX.1-2024 XCU 2.9.1-2.9.4, 2.6.3) is resolved past reserved words,
assignments, wrappers (timeout, env, nice, stdbuf, command, exec, time,
nohup, sudo 1.9.15p5, xargs), package runners and rtk proxy to a lane
executable by exact basename. mcporter operations and a call's
downstream server follow openclaw/mcporter@93e0916c (cli.ts,
command-inference.ts, call-arguments.ts, call-command.ts); HTTP and
stdio selectors read (http) and (stdio), and no host is emitted. rtk
proxy follows rtk-ai/rtk@1d87b8e7 src/main.rs:68-90,708-713,3008-3042
and the installed rtk 0.50.0 (no shell runs the proxied program).

One memoized reading per call drives carrierOf (M3 and M4 by_carrier),
measurement.proxy (adds invocations, nested, in_ctx_code and
prefix_rule_calls) and the new measurement.cli_lanes with per-call
states (AA-PLAN "Successful" and M14; not_executed from observed
transcript markers). aggregateMeasurements sums the fields and adds
actors_with_success; the sweep limits gain the static limits.
executedTrace takes an optional marks collector for ssh-run text and
data-heredoc substitutions; its output is unchanged (0 differences in
executedText and fetchKind over 24,489 inputs).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Read mcporter URL and stdio selector forms with char scanners

Replace the character-class regexes in httpUrl and httpToolSelector,
the path-prefix alternation for an ad-hoc stdio selector and the
blank split of an inline npx server with character checks and the
existing shellWords splitter. The brief asks for linear character
scanners rather than regexes in this kernel (CodeQL js/redos flagged it
before). Behaviour is unchanged: the rules still follow
openclaw/mcporter@93e0916c src/cli/http-utils.ts:1-69,
call-argument-values.ts:68-83 and ephemeral-target.ts:134-158, and the
node suite passes unchanged (107/107).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Document CLI lanes, call states and the command-position proxy rule

Add a CLI-lanes subsection to the PR-A measurement fields: the lane
executables with the gcm and serena-hooks exclusions, the resolver
(wrappers, runners, rtk proxy), each cli_lanes field, the mcporter
operation and downstream-server precedence of mcporter v0.14.1, the
call states with the observed not-executed markers and their M14
mapping, the static limits, and the changed meanings of proxy.calls,
M3 by_carrier.rtk_proxy and m4.by_carrier with prefix_rule_calls kept
for #369-era comparison. Record the decision that proxy.in_ctx_code
stays outside M6 (Gate A plan M6 and rtk rows) and state that
rtk_parts, proxy_parts and model_typed keep their prefix rules. The
skill-usage README points Codex readers to the same fields and states
that declined items and legacy exec headers still read succeeded until
the adapter maps them.

Sources: rtk-ai/rtk v0.50.0 src/main.rs:68-90,3008-3042; openclaw/mcporter
v0.14.1 src/cli and docs/call-heuristic.md; mksglu/context-mode v1.0.169
src/exit-classify.ts:15-33; POSIX.1-2024 XCU 2.9.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add a failing case: exec -- reads as an unresolved program

U1 design 2.2 reads only the POSIX exec form and marks any word that
starts with '-' after exec as unresolved, because bash's exec options
are not pinned. The shells disagree on `exec -- cmd`: GNU bash 5.2.21
runs cmd, while dash tries to execute "--" (probe on this host). The
kernel currently takes `--` as the end of options and reads qmd.
Observed before the fix: "exec -- qmd mcp" => ["qmd/qmd"].

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Read any dash word after exec as an unresolved program

Follow U1 design 2.2: only the POSIX exec form is read, so a word that
starts with '-' after exec, `--` included, leaves the program
unresolved. bash 5.2.21 runs the command after `exec --` while dash
executes "--" itself, so neither reading is safe to assume. The failing
case from 1fcc9f23 now passes (node suite 107/107).

Also cover the inline npx server of an mcporter call
(openclaw/mcporter@93e0916c src/cli/ephemeral-target.ts:134-158): a
--server value that is an npx command line is an ad-hoc stdio server,
and another blank-holding name reads (other). These cases were added
after the implementation, so they are coverage, not failing-first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add failing-first Codex fixtures for shell call states

U1 design 6 fixtures (i)-(vii) run rollout records through
measure_codex_records, the Node bridge and the kernel's cli_lanes and
proxy. Record shapes follow openai/codex rust-v0.157.1 (36650394):
CommandExecutionItem and its snake_case status (protocol/src/items.rs
:199-285), exec end states (core/src/tools/events.rs:529-573: a
rejection is declined with exit -1) and the unified exec response
header (core/src/tools/context.rs:524-575). A rollout output never
carries the success flag (protocol/src/models.rs:2173-2182).

Before the adapter change 6 of 8 subtests fail:
- (iii) code-mode item declined: qmd succeeded 1, expected failed 1
  and not_executed 1; the same for a direct call's declined item;
- (iv) legacy exec_command with "Process exited with code 1":
  rtk_proxy succeeded 1, expected failed 1;
- (v) "Process running with session ID 3": qmd succeeded 1, expected
  unknown 1;
- (vi) no header and no item: qmd succeeded 1, expected unknown 1;
- (vii) local_shell_call running stack.json:280:
  codebase-memory-mcp succeeded 1 and downstream codebase-memory
  succeeded 1, expected unknown 1 in both.
(i) nested rtk proxy (proxy.calls 1, proxy.nested 1, succeeded 1) and
(ii) a failed item already pass. Must-stay controls pass: an exit 0
header, an exit line inside the output, a completed item that beats
a running header, a failed item with its header, and the aggregate
passing cli_lanes and proxy through with actors_with_success.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Map declined items and exec headers to Codex call states

U1 design 6: the Codex adapter now hands the kernel's callState what a
rollout shows about a shell call's outcome, so cli_lanes stops reading
declined or failed legacy commands as succeeded.

- A declined CommandExecution (openai/codex rust-v0.157.1 36650394
  core/src/tools/events.rs:562-573, exit -1) sets is_error and
  native_state "declined" on both result paths: the call's own
  function_call_output and the item's aggregated output. The kernel
  counts it failed and not_executed.
- A Bash-mapped call (exec_command, shell_command, shell or
  local_shell_call, chosen by call kind, never by output text) with no
  persisted item state is read by the unified exec header of
  core/src/tools/context.rs:524-548, and only by the lines before
  "Output:", in a linear line scan with a strict ASCII exit code.
  Exit 0 is success, another exit is_error. A running line (it wins
  over an exit line, unified_exec/process_manager.rs:1066-1071), a
  header with neither line or text without the header sets
  native_state "unknown": a rollout output never carries success
  (protocol/src/models.rs:2173-2182), and legacy history mode keeps no
  item (rollout/src/policy.rs:94-112). No shell, shell_command or
  local-shell handler exists at that revision, so those outputs read
  unknown unless an item state decides; apply_patch's "Exit code: N"
  text is not read.
- codex_call_name shares the MCP namespace rule between the call loop
  and the shell-call set, so an MCP tool named exec_command is not a
  shell call.

The failing-first fixtures from 4a7aa75a now pass (the three Python
suites: 160 tests OK). The header unit test
(test_exec_header_state_reads_only_the_pinned_header) was written
after the parser and is coverage, not failing-first. The lanes CLI
still writes nothing to stderr, and the Codex aggregate passes
cli_lanes and proxy through unchanged.
tools/skill-usage/README.md replaces the stage B sentence that said
these calls still read succeeded.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Regenerate workflow SHA256SUMS for the U1 kernel and tests

The N1 fix and the CLI-lane reading changed child-usage.mjs and
test-child-usage.mjs; stage C leaves the kernel byte-identical to
0c421c66 (sha256 acc7bb51...). Regenerated with the documented
command, `sha256sum -- *.mjs *.js *.json > SHA256SUMS` in
examples/claude-native/workflows, and `sha256sum --check --strict
SHA256SUMS` passes for all 13 files. Only those two digests change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Re-register the U1 files in manifests/evidence.json

The hot-file protocol (docs/lanes.md:94-103) puts the shared manifest
edit in the branch's last commit and re-registers every changed file
that the manifest already lists, with host_receipts.register_file.
scripts/validate.py checks SHA-256 and byte counts for each of them
and failed before this commit ("SHA-256 mismatch", "byte count
mismatch" for the workflows README, child-usage.mjs,
test-child-usage.mjs, tests/test_token_measurement.py and
tools/skill-usage/README.md). The eight re-registered entries are the
two READMEs, SHA256SUMS, child-usage.mjs, test-child-usage.mjs,
skill_usage.py, test_skill_usage.py and test_token_measurement.py.
Only their sha256 and bytes change, and validate.py now passes
(status passed, 7986 hashed files).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Add failing-first tests for the U1 pivot loader and M4 repairs

Tests only, run against 7091a300 before any fix (U1 pivot brief D1, D6-D9,
D11; GPT-6 findings #8, #9, #11, #12, #13; Claude review R1, R2, R3, R9).

- D1 loader (test-child-usage.mjs, test_skill_usage.py): pin file contents,
  loadShellParser statuses (installed, not_installed, hash_mismatch),
  directory order, --shell-parser, commandInvocations null and cli_lanes
  parser_unavailable with proxy prefix_fallback, aggregate statuses,
  withShellTree freeing its tree, and the Node bridge awaiting the loader.
- D6: shell_script quotes argv with shlex.join, and -c is a script only for
  a shell; exec_header_state bounds the exit code to 9 digits (GPT-6 #11,
  #12 inputs verbatim).
- D7: outer-shell substitutions in a double-quoted run string are executed
  text (GPT-6 #8 verbatim), with single-quote, comment, backquote, sh, eval
  and ssh variants and must-stay controls that count once, not once per view.
- D8: n = 8000..64000 doubling for unclosed ((, $((, "$( ((, and ( << E
  runs, at most 2.5x per doubling and under 150 ms at 64000.
- D9: the stress check lets an exception fail it, includes an unquoted
  '$('.repeat(3000) input (NESTING_LIMIT), and a mutation control (a reader
  that throws for long input) shows the check now fails.
- R1, R3: a heredoc after a closed "$( )" keeps its reader; an escaped blank,
  semicolon or newline before a # is not a comment start.

Every expected value of the D7, R1 and R3 commands was run under GNU bash
5.2.21 and dash with a stub curl (and a stub ssh that runs sh on its stdin,
the remote login shell): both shells agreed on every count.

At 7091a300: node suite 115 passed, 31 failed (5 linear, 26 parser); D7 9
commands, R1 5, R3 3 fail with 0 != 1 on all three carriers; exec_header_state
raises ValueError on 5000 digits; shell_script flattens argv. The old D9
assertion passes when 7 of 7 stress inputs throw.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Quote Codex argv elements and bound the exec header exit code

U1 pivot brief D6 (GPT-6 findings #11 and #12).

shell_script joined an argv array with spaces, so a metacharacter inside one
element became shell syntax that was never executed: ["echo", "qmd; rtk proxy
qmd status"] counted one qmd call and one rtk proxy call. It now joins with
shlex.join (Python 3 shlex.join, the inverse of shlex.split), so each element
is one quoted word. The [shell, -lc, script] form (openai/codex rust-v0.157.1
codex-rs/core/src/shell.rs) stays the script itself, but only when the program
is a shell: -c of `echo -c ...` is one of echo's arguments.

exec_header_state passed the exit code text to int() whenever it was ASCII
digits, and CPython raises ValueError above 4,300 digits (int max str digits,
docs.python.org/3/library/stdtypes.html#int-max-str-digits), which stopped the
scan on one malformed header. The code is an i32 written in decimal (at most 10
digits with its sign, codex-rs/core/src/tools/context.rs:534-540); a code of
more than 9 digits is no native header and reads (False, "unknown").

The D6 tests of the previous commit pass: exec_header_state on 5000 digits,
the shell_script cases, and the three Codex bridge cases (local shell call,
code-mode item, `echo -c`). The existing Codex state tests still pass.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Scan unclosed (( in linear time: one paren table per text (D8, R9)

U1 pivot brief D8 (GPT-6 finding #9; Claude review R9). Every unclosed "((" or
"$((" made step() look ahead to the end of its line (indexOf the newline, then
count parentheses), so n = 8000, 16000 and 32000 of "((" took 246, 1218 and
4888 ms at 7091a300 (four times per doubling), and openers() filtered its whole
cut list once per heredoc operator.

parenClose(s) now matches every "(" of a text with its ")" in one stack pass per
line (bash(1) ARITHMETIC EVALUATION: (( )) and $(( )) count every ( and ) up to
the end of the line, which is classic parenthesis matching), and step() reads
its answer from the table. The last eight texts stay memoized so the scans that
alternate between a text and its slices do not rebuild it; executedTrace clears
the memo when a top-level analysis ends. openers() finds each operator's bounds
in the ascending cut list by binary search instead of filter and find.

Output is unchanged: executedText (both modes) and fetchKind of the base
kernel and this one agree on all 150,000 checks (the 243 covering commands, the
10,000 heredoc fuzz strings, and two seeded 20,000-string token corpora with
and without heredocs; 0 differences, 0 throws). Timings at n = 8000, 16000,
32000, 64000: "((" 3/6/13/27 ms, "$((" 4/9/19/40 ms, unclosed (( inside "$( "
5/8/18/38 ms, "(<<E" runs 8/14/30/64 ms, each within 2.5x per doubling and
under 150 ms at 64000. The D8 test allows 5 ms of timer noise on the 2.5x
bound (a linear scan showed 7.6 -> 19.4 ms once) and takes the best of five.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Read a # after an escaped blank or metacharacter as data (R3)

Claude review R3 of 3cb7c4f6. In a "$( )" frame step() treated a # as a comment
start whenever the previous character was a word break, ignoring whether that
break was escaped. `x="$(echo a\ #b)"; curl ...` therefore read `#b)"; curl ...`
as a comment: closeQuote and matchParen ran to the end of the command and the
`; curl` fell inside double-quoted data (shell_fetch 0, fetchKind null, one
unconfirmed mention). bash(1) COMMENTS: a word beginning with # starts a
comment, and an escaped blank, `;`, `(` ... or newline (a line continuation)
does not end a word (QUOTING: a backslash escapes the next character).

escapedAt(s, k) counts the backslashes before s[k]; an odd count means it is
escaped. It runs only for a # that follows a word break, and reads one run of
backslashes per such #, so the scan stays linear.

Verified with GNU bash 5.2.21 and dash and a stub curl: for each of the six
R3 inputs (escaped blank, escaped semicolon, backslash-newline, an even
backslash count before a real comment, a real comment inside "$( )", and a
top-level escaped blank) both shells ran the stub curl once, so each is one
executed fetch.

R3 test passes on all three shell carriers. Differential of the previous
commit's kernel and this one over 90,243 inputs (243 covering commands,
10,000 heredoc fuzz strings, and two seeded 40,000-string token corpora, with
and without heredocs; executedText in both modes and fetchKind): 228 differing
inputs, every one containing an escaped word break before a # (0 unexplained).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add failing-first R1 shapes and the --shell-parser refusal test

Tests only. Extends the R1 cases of 62cb6935 with the shapes the fix has to
cover beyond the reviewer's repro, each run under GNU bash 5.2.21 and dash with
a stub curl (both shells agreed):

- a quote inside the substitution: FOO="$(echo "a b")" bash <<'EOF' and the
  ssh variant, where the naive word split cuts the quoted word at the inner
  blank (curl ran once; with cat as the reader it ran 0 times);
- a string that runs over lines before the operator: x="a\nb" bash <<'EOF' and
  the single-quoted variant, and a "$( )" that closes on a later line with text
  after its ")" (curl ran once each);
- two closed "$( )" words, and controls that already read correctly at
  f83f8367 (a separator before the operator, a backquote span, a nested
  "$( )" with its own heredoc, cat as the reader).

At f83f8367 the 11 executed shapes each fail on all three shell carriers (33
subtests: shell_fetch 0 != 1), the four controls pass. Also: an explicit
--shell-parser that names a directory with no install exits 2 with the reason
and no report (the --rtk-db and --exceptions convention), never a silent
fallback to the default.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Bound each heredoc by the cuts of its own frame level (R1)

Claude review R1 of 3cb7c4f6. openers() kept one cut list for every nesting
level, so the ( and ) that a double-quoted "$( )" adds also bounded the heredoc
operators outside it: in FOO="$(pwd)" bash <<'EOF' the operator's command
started at the closing quote, its reader resolved to the empty string, and a
body that bash reads as a script became data (shell_fetch 0, one unconfirmed
mention, fetchKind null; real bash and dash run the stub curl once).

step() now reports each cut with its frame (on.cut(i, frame): the "$(" frame
that a ( opens and a ) closes, else the frame the unit sits in, undefined for
the top level) and openers() keeps one ascending cut list per level; an
operator is bounded by the cuts of the level on top of the stack, found by
binary search. The cuts of a substitution bound the commands inside it and no
longer the command that contains it. commandWords() reads a "$( )" or a
backquoted span inside double quotes as one unit of its word (POSIX.1-2024 XCU
2.6.3: the substitution's text is tokenized on its own), so a quote or blank
inside it does not split the word: FOO="$(echo "a b")" bash <<'EOF' names bash.

R1 tests (62cb6935, 58502c60) pass on all three shell carriers, and the lane
reading of FOO="$(pwd)" bash <<'EOF' with qmd in the body counts one qmd call.
The shapes below are documented limits, pinned by their own test: a command that
a quote or "$( )" carries over lines (x="$(cat <<'A' ... \n)" bash <<'B'
and x="a\nb" bash <<'EOF') has its head on an earlier line, so this line alone
gives the operator no reader and the body reads as data (fetch_mentions_unconfirmed
1, status incomplete: a possible fetch, never a lost one). A first attempt cut
the line at the closing quote's word end; measured on the fuzz corpus it named
the wrong owner as often as the right one (ssh host cat "a\nb" bash <<EOF), so
it was dropped.

Differential of f83f8367 and this commit over 90,243 inputs (243 covering
commands, 10,000 heredoc fuzz strings, two seeded 40,000-string token corpora,
without and with heredocs; executedText in both modes, fetchKind and the M4
confirmed and unconfirmed counts): 17 inputs differ, all with a quote,
backquote or "$( )" before a heredoc operator (0 unexplained); no input lost
mass (confirmed plus unconfirmed); 16 keep both counts, one moves a fetch from
confirmed to unconfirmed. That one, bash -s -- "$(pwd)"bash <<EOF with an
unquoted-delimiter body, now names bash as the reader (correct) and reads the
body as source, where a substitution inside single quotes is data: the
"$( )"-free twin reads C=0 U=2 at 7091a300 as well (real bash runs curl once:
the outer shell expands an unquoted heredoc body before the reader sees it, a
limit that predates R1 and is unchanged).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Read a double-quoted run string in two views: outer expansion (D7)

U1 pivot D7, GPT-6 finding #8. N1 read a double-quoted string that a shell runs
only as that shell's input, so a substitution the OUTER shell runs while it
expands the word was lost whenever the inner shell reads it as data: a quoted
heredoc, single quotes, a comment. bash -c "cat <<'EOF'\n$(curl https://example.org)\nEOF"
went from fetchKind 'fetch' (cf3fb72e) to null with shell_fetch 0. bash 5.2.21 and
dash run the stub curl once: the expanding shell runs each unescaped "$( )" and
backquote before the shell it starts reads anything (POSIX.1-2024 XCU 2.2.3
Double-Quotes, 2.6.3 Command Substitution). Two views now: outer bodies first,
each as commands of their own, then the inner shell's read of the string with
each outer span replaced by a placeholder word (its output is unknown), so a
substitution both views could see counts once. A backslash-escaped \$( ) or
backquote is no outer span and stays the inner shell's (bash -c "echo \$(curl u)"
runs curl once, in the inner shell). Single-quoted run strings have no outer
expansion and are unchanged.

outerSpans() is the one scanner of a double-quoted string's substitutions (escape
pairs skipped, "$((" arithmetic not a span but the substitutions inside found,
NESTING_LIMIT and unterminated spans as before); quotedData() now consumes it, so
the data view and the run view cannot disagree. The ssh remote ranges (marks)
cover the literal text between spans, not the local substitutions.

D7 tests (62cb6935) pass on all three carriers: 13 executed, 5 inner and 2
twice-run commands, 2 data commands; the qmd lane test counts one call.

Evidence, all against GNU bash 5.2.21 and dash with stub curl/wget/ssh (ssh runs
its arguments with sh -c, or reads stdin as the remote login shell), counting
stub calls against the kernel's confirmed count (shell_fetch + loopback):
- compositional oracle, 360 commands (4 double-quoted run carriers x 10 contexts
  x 4 atoms, two-atom strings, single-quoted carriers, heredoc readers x
  delimiters x bodies): base 7091a300 301 exact, 59 under, 0 over; 43b9d947 329,
  31, 0; this commit 353, 7, 0. The 7 left are one class, a shell-read heredoc
  with an unquoted delimiter whose body has a substitution inside single quotes
  (the outer shell expands the body first), next commit.
- random valid-bash fuzz, 40,000 candidates, 7,883 accepted by bash -n and dash -n
  with equal stub-call counts under both (361 with a fetch): base, f83f8367,
  43b9d947 and this commit agree on the same 7,789 (98.81%) and the same 94
  disagreements (12 over: word glue and a trailing backslash; 82 under: remote
  ssh commands and glued words), all present at 7091a300.
- differential of 43b9d947 and this commit over the covering commands and three
  fuzz corpora (90,243 inputs): 13 inputs lose confirmed mass, every one invalid
  bash in at least one shell; of the 24 differing inputs without HTTP-library or
  gh tokens, 21 are invalid, 2 valid inputs move closer to real bash and 1
  (xcat &&bash -c "'$(curl u)'") counts a substitution that a failing && never
  reaches, which no static reading models.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add failing-first tests: an unquoted heredoc body is expanded first

Tests only. D7 for a heredoc body (U1 pivot D7, same defect class as GPT-6 #8):
with an unquoted delimiter the shell that reads the command line expands the
body before the shell that reads the heredoc sees it, so each unescaped
substitution runs in the outer shell whatever its quotes or a comment say
(bash(1) Here Documents; POSIX.1-2024 XCU 2.7.4), quotes are literal, and a
backslash escapes only $ ` \ and a newline. 20 commands, each run under GNU bash
5.2.21 and dash with stub curl and ssh: curl ran once for 17 (a single-quoted
substitution, a comment, nested quotes, a substitution over lines, a <<- body
with tab indents, arithmetic, an ssh reader, a run string in the body, and
controls that already read correctly), twice for one (an outer and an inner
substitution), and never for two (a quoted delimiter passes the body as is).

At 67a4acbb 9 of the 17 fail on all three shell carriers (shell_fetch 0 != 1,
one unconfirmed mention: bash <<EOF with echo '$(curl u)', the FOO="$(pwd)" and
ssh readers, a comment, nested quotes, a multi-line substitution, <<-, and
arithmetic); the rest and the data commands pass.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Expand an unquoted heredoc body a shell reads before reading it

The heredoc form of D7 (GPT-6 #8: the outer shell runs what it expands before
the inner shell sees the text). With an unquoted delimiter the shell that reads
the command line expands the body first: each unescaped $( ) and backquote runs
there whatever its quotes or a comment say, quotes are literal in a body, and a
backslash escapes only $ ` \ and a newline (bash(1) Here Documents; POSIX.1-2024
XCU 2.7.4). resolveHeredocs() analyzed a body that a shell (or in inlineHttp
mode an interpreter) reads only as that reader's source, so bash <<EOF with
echo '$(curl u)' read as data although real bash and dash run curl once, and
so did a substitution in a comment, nested quotes inside it, one over lines,
arithmetic, and a <<- body: the same 9 shapes as the failing-first commit
b6f3b454. R1 made the FOO="$(pwd)" bash <<EOF form read this way too, so this
also removes the accident that had hidden the limit there.

The body now has two views, sharing D7's outerSpans/withoutSpans/outerBody: the
outer spans first, as commands of their own (heredocs inside them resolved
there), then the reader's source with each span replaced by the placeholder word
and its own escapes applied (heredocText). A quoted delimiter is unchanged
(nothing is expanded), and so is a data heredoc such as cat <<EOF, which keeps
the substitutions its regular expression finds. The ssh remote ranges cover the
body's literal text, not its local substitutions.

The heredoc test of b6f3b454 passes on all three carriers (17 once, 1 twice, 2
data commands); tests.test_token_measurement passes (59). Evidence, real bash
5.2.21 and dash with stub curl/wget/ssh, kernel confirmed count against stub calls:
- compositional oracle, 441 commands (D7's 360 plus escaped, comment and
  nested-quote bodies): base 7091a300 368 exact, 73 under; 67a4acbb 420, 21;
  this commit 441 of 441 exact, 0 under, 0 over.
- grammar fuzz, 6,000 valid nested commands (statements, sequences, substitutions,
  backquotes, single and double quotes, run strings, heredocs with unique
  delimiters) that bash and dash ran with equal stub-call counts (3,049 with a
  fetch, 1,597 with a run string, 1,471 with a heredoc): base 5,788 exact
  (96.47%), 212 under, 0 over; 67a4acbb 5,970 (99.50%), 30 under, 0 over; this
  commit 5,975 (99.58%), 25 under, 0 over. All 25 are one older gap, a run string
  right after a backquote (echo `bash -c "curl u"`), whose text is invisible to
  the raw detector as well; fixed next.
- random valid-bash fuzz, two seeds of 40,000 candidates (7,883 and 7,838 valid,
  361 and 374 with a fetch): base, 67a4acbb and this commit agree on every
  input.
- differential of 67a4acbb and this commit over 90,243 inputs: 924 differ, 20 with
  a lower confirmed count, of which 19 are invalid bash or valid only with a
  heredoc-at-end-of-file warning, and the one clean input (a backquote span that
  holds bash <<EOF) moves from 3 to 0 confirmed where real bash and dash run
  curl 0 times.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add failing-first tests: a run string after an opening backquote

Tests only. RUN_QUOTED anchors a shell, eval or ssh word at the text start, a
blank, ; & | or (, but not at a backquote, so bash -c "curl u" inside `...` read
as double-quoted data. The fetch was lost, and the raw detector, which anchors
curl itself and not the quote, counted no possible fetch either (C=0 U=0). Found
by the grammar fuzz of 2a33ef90: all 25 commands left under-counted there were
this one shape.

8 commands run under GNU bash 5.2.21 and dash with stub curl and wget: the stub
ran once for each (bash -c and sh -c with double or single quotes, eval, an ssh
string, a backquote inside a single-quoted run string, an escaped backquote in a
double-quoted one, and two controls that already read correctly: a blank after the
backquote and $( ) in its place) and never for echo `echo "curl u"` (data).

At 2a33ef90 the 6 backquote-adjacent shapes fail on all three shell carriers
(shell_fetch 0 != 1, no unconfirmed mention: 19 failures with fetchKind).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Anchor a run string at an opening backquote too

RUN_QUOTED and RUN_HTTP_CODE anchored a shell, eval, ssh or interpreter word at the
text start, a blank, ; & | or (, but not at the backquote that opens a `...`
substitution, so bash -c "curl u" inside backquotes read as double-quoted data: the
fetch was lost and no possible fetch was counted either (C=0 U=0). bash(1)
Command Substitution: a `...` body is shell text with its own command positions,
as a $( ) body is; the anchor class now holds the backquote (one character in
each regular expression; no new alternative, so no new backtracking).

The 6 backquote shapes of the failing-first commit (66550644) pass on all three
carriers; tests.test_token_measurement passes (60).

Evidence, real bash 5.2.21 and dash with stub curl/wget/ssh, kernel confirmed count
against stub calls: the 441-command oracle stays 441 exact; the grammar fuzz that
found the gap is now exact on every input, 6,000 of 6,000 (3,049 with a fetch,
1,597 with a run string, 1,471 with a heredoc; 2a33ef90 5,975 exact, 7091a300
5,788, 96.47%) and, on a second seed, 11,999 of 11,999 (base 11,608, 96.74%, 391
under; 0 over throughout). Differential of 2a33ef90 and this commit over 90,243
inputs: 216 differ, 1 with a lower confirmed count, an unterminated string (invalid
bash).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add a pinned, verified tree-sitter-bash loader that fails closed (D1)

U1 pivot D1. The CLI-lane reading is to run on tree-sitter-bash, the grammar
OpenAI Codex uses for the same job (openai/codex rust-v0.157.1 36650394
codex-rs/shell-command/src/bash.rs; tree-sitter-bash = "0.25" at
codex-rs/Cargo.toml:522, both read at the tag). This commit adds the seam and the
gate; the AST reading itself is the next stage.

shell-parser.pin.json pins the install: web-tree-sitter 0.27.0 and
tree-sitter-bash 0.25.1 with their npm integrity values (checked against the
registry's dist.integrity), upstream tags and commits (v0.27.0 6070dbfe =
npm gitHead; v0.25.1 names a06c2e44 while the package was published from
80132668, five minutes earlier, recorded as a note), the sha256 of the six
files the loader reads, and the install command (npm install --prefix <dir>
--ignore-scripts --no-audit --no-fund --save-exact). loadShellParser(dir)
resolves the directory as the argument (the new --shell-parser flag), then
CHILD_USAGE_SHELL_PARSER, then the ecosystem tools directory under HOME, reads
each pinned file once, compares its sha256 and the lockfile's version and
integrity values with the pin, and executes only those bytes: the runtime module
is imported from a private 0600 copy of the hashed bytes and the two wasm files
are handed over as bytes, so the code that was hashed is the code that runs.
Results carry no path: { ok, versions, wasm_sha256 } or { ok: false, reason }
with not_installed (directory, a pinned file or the lockfile missing),
hash_mismatch, load_error (an unreadable file, the pin file, init) and
not_loaded (this process never awaited the loader, so a forgotten await is not
mistaken for a missing install). A failed load leaves no parser, whatever an
earlier call did; a successful one reuses the initialized runtime. withShellTree()
frees every tree in a finally.

Fail closed: commandInvocations() returns null without a parser and measurement
reports cli_lanes { status: parser_unavailable, reason }, keeps the prefix rule for
the rtk proxy carrier (proxy.rule prefix_fallback, invocations and in_ctx_code
null) and never falls back to the text scanners for lane counting; M4 stays text
based. A measured cli_lanes carries status measured and the parser record (versions
and the two wasm sha256 values); aggregates report measured, incomplete (mixed
actors, with the counts), parser_unavailable or not_measured, and sum proxy
counters null-aware. The CLI awaits the loader; a --shell-parser that cannot be
honored exits 2 with the reason (like --rtk-db), any other missing install only
leaves cli_lanes unavailable. The Codex bridge awaits the loader too.

Interim, until the AST reading lands: with a verified parser loaded,
commandInvocations() still runs the scanner-based lane layer of 3cb7c4f6, so every
existing lane fixture reads as before; the gate, the status fields and the seam
(withShellTree) are final.

Tests (62cb6935, 58502c60 first, run failing at 7091a300): the pin's contents;
not_installed for an absent directory, a missing pinned file and a missing
lockfile; hash_mismatch for one flipped byte in the bash wasm, a modified
web-tree-sitter.js that is never imported, a modified package.json, another
integrity value and another version in the lockfile; load_error for an unreadable
file; directory order; --shell-parser; a good install loading again after
failures; unavailable and mixed aggregates; the bridge in both states. The lane
fixtures load the parser first and skip with a message on a host without an install.
tests: node suite 149 passed; tests.test_token_measurement, test_skill_usage and
test_child_usage_suite 175 passed. D9's stress check reports the TypeError the
lane checks threw before the parser was loaded (threw: TypeError x8), which the
old assertion would have passed.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add failing-first timing tests for quote-heavy scripts (D8)

Tests only. D8 asks the M4 text scanners for linear time; the unclosed "((" input
of GPT-6 #9 is fixed (5e814331), but profiling a realistic 346 KB script (4,000
lines of echo "step N: $(date)" >> log; qmd search "term N" | head) at e5238f39
showed the same quadratic shape from another term: 51% of the time in the
run-string regular expression and 33% in scanQuotes, both from one read of the
whole text built so far per quote (the regex, and the string flattening that
indexing a concatenation rope forces). A 42 KB script takes 47 ms, 84 KB 108 ms,
172 KB 467 ms, 346 KB 2,078 ms and 694 KB 8,236 ms a scan, and the kernel scans
each shell call several times.

Five doubling checks (n = 8000..64000, best of five, at most 2.5x per doubling
plus 5 ms of noise, 64,000 under 1.5 s): a run of quoted words, a run of
double-quoted substitutions in inlineHttp mode, a run of `bash -c "x"` strings,
a run of words with a # inside, and a long script of echo, substitution and pipe
lines through fetchKind. At e5238f39: 61, 230, 902, 5,861 ms; 168, 651, 3,728 ms;
991, 4,486 ms; 9, 32, 120, 459 ms; 89, 392, 1,930 ms (the series end where a run
passes 1.5 s). All five fail.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Read only the tail of the built text at each quote (D8: linear scripts)

scanQuotes() tested RUN_QUOTED (and in inlineHttp mode RUN_HTTP_CODE) against the
whole text built so far at every quote. Reading out.s for that flattens the string
concatenation and scans the full prefix each time, so a script with n quoted
words cost n squared. CPU profile of a 346 KB script (4,000 lines of echo "step N:
$(date)" >> log; qmd search "term N" | head) at e5238f39: 51% in the run-string
regular expression, 33% in scanQuotes itself (the flattening), 10% in GC. The
five timing tests of 9da57414 failed on it.

scanQuotes now keeps the end of its output in `tail`, cut back to 512 characters
whenever it passes 1,024, updated at every write (put and putTrace; quotedData
builds into its own buffer first), and reads only that: the # test uses the last
character of the tail, the emptiness test out.p.length. Once a cut has been made
the two detectors run without their start-of-text alternative (a cut is no start
of text). Limit, stated in the code: a run keyword more than 512 characters
before its quote, after a text that has grown past 1,024, is missed and the
string reads as data. Only a single ssh option word that long can put it there
(ssh -oProxyCommand=<1500 characters> host "curl u": fetch before, null after;
100, 400, 520 and 700 characters read the same), since the words between a
shell, eval or ssh and its string are short options and one host.

Timings (best of five, ms), n = 8000, 16000, 32000, 64000: quoted words 61, 230,
902, 5,861 -> 8, 15, 31, 69; double-quoted substitutions (inlineHttp) 168, 651,
3,728 -> 15, 31, 67, 140; run strings 991, 4,486 -> 38, 83, 173, 358; words with
# 9, 32, 120, 459 -> 3, 6, 14, 37; a long echo/substitution/pipe script through
fetchKind 89, 392, 1,930 -> 18, 34, 71, 154. Realistic scripts, executedText: 84
KB 108 -> 37 ms, 346 KB 2,078 -> 100 ms, 694 KB 8,236 -> 205 ms, 1.4 MB 393 ms
(2x per doubling); the node suite runs in 12 s instead of 41 s.

Output is unchanged: executedText (both modes) and fetchKind of e5238f39 and this
commit agree on all 1,084 shell texts of this repository (822 .sh files and fenced
shell blocks of .md files, 1.66 MB, largest 82 KB, 83 over 4,000 characters), on
the 243 covering commands and the three fuzz corpora (90,000 inputs), and on 600
synthetic long scripts (0 differences, 0 throws); the 441-command oracle stays 441
exact, and the 6,000-command grammar fuzz stays 6,000 exact (real bash 5.2.21 and
dash with stub curl).

Also fixed while rewriting the call: the ssh remote offset test listed the
separators before RUN_QUOTED's keyword but not the backquote that 296c4f1d added,
so in echo `ssh host "qmd get a"` the qmd call read as local; it now reads
qmd/qmd remote (fixture added; the same input read local at e5238f39).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Flip one real byte in each pinned file in the loader tests

The loader tests changed the bash grammar wasm by reading it as UTF-8 and
appending a NUL, which rewrites the whole binary: the test passed because the
file differed, not because one byte did. It now reads a Buffer and flips bit 0 of
the middle byte in place (same length, so only the hash can tell), and does the
same for each of the six pinned files in turn: every one gives hash_mismatch,
leaves no parser (commandInvocations returns null, the scanners do not step in) and
runs no unverified code (the modified web-tree-sitter.js sets a global if it is
ever executed; it is not). The missing-file and missing-lockfile cases now also
assert that no parser stays loaded. Node suite: 160 passed (149 before).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add a failing-first test: a run keyword past the tail window

Tests only. The brief's D8 says a work budget "marks the command unclassifiable for
M4 instead of throwing". 1d77c34c bounds scanQuotes to the last 512 characters of
its output, but an ssh option word long enough to push `ssh` out of the window
(in a text past 1,024 characters) drops the run string silently: the fetch in
ssh -oProxyCommand=<1500 characters> host "curl u" is read as data with no
possible fetch counted either (shell_fetch 0, unclassifiable 0, unconfirmed 0),
which a budget must not do.

The test: option words of 100, 400 and 700 characters read as before (the fetch
confirmed, unclassifiable 0); 1500 characters counts one unclassifiable operation
(remote_fetches 1, no confirmed shell fetch); the count is once per command and
only for the command that held the keyword (two later quoted commands add nothing);
a long printf with hundreds of words and quotes but no run keyword adds none.

At 1d77c34c the two 1500-character cases fail on all three shell carriers
(unclassifiable 0 != 1: 6 failures); the 100, 400 and 700 cases and the printf
control pass.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Count a run string the tail window lost track of as unclassifiable

The brief's D8 fallback for input past a work budget is "marks the command
unclassifiable for M4", not a silent miss. 1d77c34c bounds scanQuotes to the last
512 characters of its output; an ssh option word long enough to push `ssh` out of
the window (in a text past 1,024 characters) dropped the run string with no trace:
ssh -oProxyCommand=<1500 characters> host "curl u" read as shell_fetch 0,
unclassifiable 0, unconfirmed 0.

scanQuotes now notes, outside quotes, when a shell, eval, ssh or su word (and its
blank) has been written since the last command separator (a last-letter pre-check
keeps it off the hot path). When a cut has left the tail without that word and
without a separator, the next quote that does not read as a run string is counted
once in windowMisses, which countFetches adds to unclassifiable (an operation of
unknown kind: remote_fetches 1, no confirmed shell fetch). The keyword is noted at
write time because the discarded text cannot say whether a word was quoted: a first
version that looked for the word in the discarded text also fired on 4 real scripts
of this repository (adoption/bootstrap-linux.sh, bootstrap-macos.sh,
launchd-agents.sh and a memory-stack README block), where separators inside quoted
data are blanked and a word such as sh sat in a long quoted list.

The test of 0b2b5778 passes on all three carriers (100, 400 and 700 characters
read as before; 1500 counts one unclassifiable, once for the command, and adds
nothing for later commands or for a long printf with no run keyword).

The counter is zero on every corpus: complete m4 results of the pre-window kernel
(e5238f39) and this one are identical on the 243 covering commands, 10,000 heredoc
fuzz strings, two 40,000-string token corpora and all 1,084 shell texts of this
repository (822 .sh files and fenced shell blocks, 1.66 MB), 2 tools each; the 600
synthetic long scripts agree; the 441-command oracle is 441 exact and the grammar
fuzz 11,999 of 11,999 exact against real bash and dash. Realistic scripts, executedText:
84 KB 50 ms, 346 KB 125 ms, 694 KB 285 ms (2,078 and 8,236 ms at e5238f39).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Keep 150 ms for the brief input only; loosen the extra timing shapes

The brief's D8 requires at most 2.5x per doubling and under 150 ms at n = 64,000 for
'(('.repeat(n) + 'qmd'. The timing tests added alongside it (a run of $((, the same
"$( ((, a run of ( << E, and the quote-heavy shapes) had inherited that absolute
bound. Three of them run 64-80 ms at 64,000 against 150 ms on this host, thin
headroom for a slower CI runner. Only the brief's input keeps 150 ms (it runs about
30 ms); every other shape keeps the 2.5x ratio (plus 5 ms of timer noise, best of
five) and takes an absolute bound of 1.5 s at 64,000, which is still far below the
quadratic behavior they replaced (the same shapes take 1.9 to 10 s and fail the ratio
at 7091a300).

Run against the base kernel (7091a300) with this test file, all ten timing shapes
fail (for example (( 313, 1,219, 5,109 ms; quoted words 62, 239, 928, 6,211 ms), and
against this head all ten pass (for example (( 3, 7, 17, 32 ms; quoted words
8, 15, 32, 71 ms; run strings 42, 95, 205, 416 ms).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Retry the doubling timing rounds so one pause cannot fail a linear scan

A direct run of test-child-usage.mjs at da2d578d failed one assertion, "a run of unclosed
(( at n = 8000, 16000, 32000, 64000 takes [4, 7, 27, 31] ms": the 32,000 reading was 27 ms
where the scan takes about 15, so the ratio to the 7 ms before it (3.9x) broke the 2.5x
(+5 ms) bound. The same suite had passed inside the Python run seconds earlier. A flaky
timing check is a CI hazard, so the doubling measurement now takes the best of five per
size and, when the ratio still fails, up to two more rounds and the elementwise minimum of
all of them: a pause cannot fail a linear scan, while a quadratic one fails every round.

Checks on the final file: six consecutive runs and two runs beside six busy CPU loops, all
160 passed; against the base kernel 7091a300 all ten timing shapes still fail (the retry
does not hide quadratic behavior).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add failing-first tests for the AST command-position layer (D10, D11)

The GPT-6 review of the scanner-based lane reading (3cb7c4f6) found ten
command-position defects (#1-#7, #10 here; #8, #9, #11-#13 were repaired
earlier) and the Claude review five more (R4-R7). This commit holds their
probes before any repair, so each one fails for its defect.

- tests/test_token_measurement.py: one test per finding, with the GPT-6
  inputs verbatim, asserting what the finding states (lane calls, proxy
  calls, exclusions, unresolved programs, downstream keys, call states),
  not a display string. Also parse_errors (D2), the closed name vocabulary
  (D5) and mcporter's help/version tokens (R6: checked against
  openclaw/mcporter@93e0916c src/cli.ts:137-145 and flag-utils.ts:34-41,
  the code already matches; the test pins it).
- test-child-usage.mjs: a program name is emitted only for a lane
  executable or a name the reading interprets (GPT-6 #10).
- tests/test_command_position_oracle.py (D10): a deterministic generator
  and a hand-written probe list run under REAL bash 5.2 with env -i, a
  PATH of logging stubs, a temporary HOME and cwd and a 5 s limit; the
  multiset of lanes that ran must equal the lanes commandInvocations reads.
  Against the scanner layer: 38 of 99 probes and 59 of 457 generated
  commands disagree with the run.

Every expected value is the run of real bash and dash under stub
executables (the oracle covers each input), not what this kernel says.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Read shell text with the tree-sitter-bash AST for CLI lanes (D2-D4)

Replaces the scanner-based command-position layer (simpleCommands,
resolveInvocation, analyzeScript) with a walk of the tree-sitter-bash
tree, the grammar openai/codex rust-v0.157.1 codex-rs/shell-command
src/bash.rs uses for the same job. The option tables, runner maps and
mcporter grammar are unchanged and now operate on word values.

- D2: every `command` node counts wherever it sits (substitutions inside
  strings, arithmetic, arrays and unquoted heredoc bodies, subshells,
  compound and control statements, function bodies); array elements, case
  patterns, [[ ]] operands, function names and the words of any non-lane
  program never do. ERROR nodes and what is under them are skipped and the
  call counts once in cli_lanes.parse_errors.
- D3: wordValue() returns a literal or null: unquoted escapes, '...', the
  double-quote rule (a backslash goes only before $ ` " \ and newline),
  $'...' decoded as bash does, concatenations; any expansion, glob, brace
  expansion is unknown. A tilde prefix keeps the word known for program
  identity (only the directory changes: ~30 lane invocations by ~/path in
  136,361 real commands) but not as a script.
- D4: bash|sh|dash|zsh|ksh -c reads the first operand after the options
  (-n, -D and -o noexec run nothing; `--`, `-` end the options); a shell
  with no -c reads standard input, so a heredoc or here-string attached to
  the last command of a list, pipeline or negation, through env, timeout,
  nice, nohup, stdbuf, rtk proxy (not xargs, which consumes it), has its
  body read; eval and ssh join their words as the callee does; rtk proxy
  resolves its program with exec semantics (command and exec run nothing,
  time is unresolved, an expansion in its one argument is unresolved).
- Backquotes and outer substitutions of an unquoted heredoc body and of a
  double-quoted script are read from the raw text with the kept M4 helpers
  (tree-sitter-bash leaves backquotes in heredoc_content and mis-reads
  some bodies); a script with an expansion is read with a placeholder.
- The grammar folds the next line into a command in some shapes (an
  unquoted == before ;, &&, | or a newline; a trailing blank before a
  newline; `<<'EOF' 2>&1 | tail`): 732 of 136,361 real commands with no
  error and 163 with one. A separator ERROR or an unescaped newline inside
  one command's words starts a new command.
- simple counts assignment-only statements, declarations and each loop,
  if and case as well as commands (POSIX.1-2024 XCU 2.9.1, 2.9.4), so the
  ambiguity of existing fixtures is unchanged.

Fixtures: EMPTY_CLI gains parse_errors (D2), and `ssh host bash -s <<EOF`
now also reads the remote bash as a command of its own (ssh(1): the words
after the destination run remotely). Every other existing lane fixture
passes unchanged.

Against real bash 5.2: the oracle agrees on 457 of 457 generated commands
(581 lane runs) and 99 of 99 probes. Over 136,361 distinct real commands
the old and new readings agree on 136,167; the differences are old false
positives, unknown wrappers (setsid, bwrap) and grammar defects. Still
failing here by design, fixed next: D5 vocabulary (5 tests) and R7 states
(7 tests); the node suite fails only on the D5 program-name check.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Re-expect the closed name vocabulary in the lane fixtures (D5)

U1 pivot D5 and GPT-6 #10: commandInvocations returns `program` only for
a lane executable or a name this reading interprets itself (a wrapper, a
shell, eval, ssh, rtk), and an mcporter downstream key is emitted only for
the stack's own servers; every other name-shaped string (a host, an id, a
package, `linear`) reads null or (other). The old expectations that showed
`-/echo`, `-/git`, `-/pytest` or `@linear` are wrong under D5, so they are
changed here, before the code: 7 node checks and 3 Python tests fail
against the walker of 1d35df12 with the observed values recorded in the
log (`program` still holds "my-private-host.example", the downstream key
still holds "linear" and "call_PRIVATE").

The selector-form coverage the `linear` cases gave is kept with servers of
the stack (`socraticode`, `serena`, `jcodemunch`, `ai-memory`, `qmd`,
`headroom`, `context-mode`, `codebase-memory`), read by name in every form.
Sources of the closed set: manifests/stack.json:280 (codebase-memory),
:407 (context-mode), :629 and :966 (socraticode); the rest is the
coordinator's list in the pivot brief.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Emit only fixed names: a closed program set and mcporter server set (D5)

GPT-6 #10: `program` carried any name-shaped basename (a host, an id, a
package) and an mcporter downstream key any name that passed a character
class, so the no-id, no-host output contract held only by luck.

- `program` is set only for a lane executable or a name the reading
  interprets itself (the wrappers, the shells, eval, ssh, rtk); every other
  program reads null.
- An mcporter downstream key is one of codebase-memory, context-mode,
  jcodemunch, serena, socraticode, qmd, headroom or ai-memory (the
  stack's servers: manifests/stack.json:280, :407, :629, :966 and the
  coordinator's list), else (other), (http), (stdio) or (unresolved).

The tests of the previous commit now pass (node 162 of 162; the Python
name-vocabulary test and the three re-expected fixtures).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Expect an interrupted counter in lane and downstream rows (R7)

Claude review R7: a Claude Code result whose toolUseResult.interrupted is
true read as succeeded. The command started and was cut short, so it is a
state of its own. The shared fixture helpers (lane_row, downstream_row and
the Codex cli_lane_row and downstream row) now carry `interrupted: 0`; the
lane and downstream rows of the kernel do not have the key yet, so the
tests that compare whole rows fail (see the log of this run) until the next
commit adds the counter and the state.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Read not-executed calls where the client records them; add interrupted (R7)

Claude review R7: the marker "The user doesn't want to proceed" was matched
against the row's toolUseResult, where it was never observed; it opens the
result CONTENT. Observed on this host's Claude Code transcripts (122,648
Bash results in 4,355 files, count-only): the rejection content (25 rows,
toolUseResult "User rejected tool use"), a row-level toolDenialKind on
every client denial (permission-rule 750, user-rejected 46, cancelled 1,
automode-unavailable 1), the classifier text "The server-side auto mode
classifier gave no verdict" (its own text says the check failed, not the
action), and 209 rows of a host hook's refusal text ("This agent is
isolated") that carry no toolDenialKind and stay failed (open: its meaning
is the hook's, not the client's).

- callState: not_executed when the row has a toolDenialKind, or the content
  opens with <tool_use_error>, "The user doesn't want to proceed" or the
  classifier text, or toolUseResult names a hook denial, a permission
  denial or a rejection.
- toolUseResult.interrupted === true is its own state, `interrupted`
  (none of 25,506 observed object results, so it is a synthetic fixture),
  with a counter in lane and downstream rows; it no longer reads succeeded.

The R7 test of the first commit and the fixtures of the previous commit
pass (measurement and Codex modules: 188 tests).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Remove the lane-only marks plumbing from the M4 text machinery

The scanner-based lane reading asked executedTrace for `marks` (the raw
range of each ssh body and the offset of each command substitution of a
data heredoc body). The AST walker reads both from the tree and the raw
text, so nothing passes marks any more: the parameter and every branch
that wrote to it (resolveHeredocs, scanQuotes, outerBody, quotedData) and
rawRange are removed. The M4 readings are unchanged: the complete m4
object of measureTranscript (Bash and ctx shell carriers), executedText in
both modes and fetchKind are identical between the kernel before the
walker (f1ed98ac) and this one over 136,361 distinct real commands, the
243 covering commands, and 2 x 40,000 fuzz inputs (0 differences, 0
exceptions).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Test the recovery of commands the grammar joined to the previous one

The walker of 1d35df12 starts a new command at a separator ERROR or an
unescaped newline inside one command's words, because tree-sitter-bash
0.25.1 reads what follows a command as its arguments in these shapes
(measured on this host's transcripts: 732 of 136,361 distinct real
commands with no error and 163 with one): an unquoted == or =~ before ;,
&&, | or a newline, a redirection after a here-document operator followed
by | or ;, and a multi-line shape found in the real corpus. This commit
adds the tests that pin it:

- tests/test_command_position_oracle.py: 12 recovery probes whose expected
  lanes are what REAL bash 5.2 ran (12 of 12 agree; 10 carry a parse
  error, 2 none).
- tests/test_token_measurement.py: the same shapes through
  measureTranscript, with parse_errors counted once for the calls whose
  tree has an ERROR, and two controls that are not boundaries (a
  backslash-newline continuation and a comment).

Mutation control (recovery disabled in a scratch copy of the kernel): all
12 oracle probes fail (real bash ran the lane, the kernel read none or
one fewer, e.g. `echo ==; qmd get a` ran qmd, read []) and 9 of the 12
measurement cases fail with the observed values ({} != {'qmd': 1}).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Add failing-first tests: a line that begins with a backslash

Found by the 14,461-command run of the oracle generator over more seeds
against real bash: a line that begins with a backslash (`\ls`, the alias
bypass; `\markitdown a.json`) reaches tree-sitter-bash 0.25.1 as a word that
begins with the newline, so the line joins the command before it, and when
it is the first line of a here-document body the grammar leaves the body
node without it and adds its words to the operator's arguments (which hid
that the owner is a shell). A here-string that precedes such a line was fed
to the last command of the joined node instead of its own.

Five shapes are added to the oracle's recovery probes and to the
measurement recovery test; real bash ran every lane in them. Against the
kernel of 907ac35f the two tests fail with the observed values (recorded in
the log of this run), e.g. `{ timeout 5 repomix` + newline +
`\markitdown --flag; }` read {'repomix': 1} where bash ran both, and the
shell heredoc whose first body line starts with a backslash read {}.

Also hardens the oracle generator: it drops a text bash rejects
(`bash -n`), keeps a function definition and its call in one group,
does not put `!` or a heredoc on the right of a pipe (`x | ! y` is not
POSIX; a reader that ignores its pipe kills the writer with SIGPIPE), and
writes data heredocs to /dev/null.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

* Read a line that begins with a backslash as a command of its own

Repairs the shapes of the previous commit's tests:

- commandSegments starts a new command at a word that begins on a new
  line (the grammar puts the newline of a `\cmd` line at the start of its
  word) and drops that newline from the word's value.
- A here-document operator's arguments are only those on its own l…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant