Skip to content

SHELL-DAG-CENSUS-0A: derive shell execution routes from parsed bodies (brief: docs/plans/shell-dag-census-0a-brief.md) - #8662

Merged
briansrls merged 4 commits into
mainfrom
session/sunny-pike-156
Aug 20, 2026
Merged

briansrls merged 4 commits into
mainfrom
session/sunny-pike-156

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor

What this is

SHELL-DAG-CENSUS-0A was dispatched to derive every authored path that can attempt POSIX/Bash program interpretation, from parsed bodies, with a hard stop condition:

If the available parsed-body facts cannot preserve exact callee identity, argument edges, branch identity, or interprocedural value flow, stop. File the missing typed projection as a separate substrate increment. Do not substitute text scanning.

The stop condition fires. This PR is docs-only: no detector, no census, no text-scanning substitute. Branched from a6ca6882d (the first main SHA meeting the brief's precondition — #8652 merged, witnesses SUCCESS on that exact SHA, run 32370279129).

Why

The whole typed body-fact surface .dag can read is three accessors in coproduct_reflection.rs — ItemKind::TypeItem, FnItem|FuncItem, DataItem. Five axes fail, each with a different remedy:

  1. Service operations and transports are absent entirely. ItemKind::ServiceItem has no accessor, so transport shell { argv: ["bash","-s"] stdin: script.body } is invisible. The brief permits exactly one seed — what interprets bytes as shell — and in this tree that fact lives only in these declarations. A census over the available surface cannot have a permitted seed at all.
  2. FnArrowDecl.output is a wiring-liveness skeleton, not a body. marshal_stmt_sequence drops a let b = rhs whose bound name is not referenced toward the return — correct for dead-wire detection, fatal here. A shell call in a discarded binding is absent from the input, so this is a structural false zero that none of the brief's seven instrument controls can catch: they all guard the detector, and the loss already happened upstream of it.
  3. No named-argument edges — labels erased, literals hoisted to the call node.
  4. No arm or occurrence identity — so the 36-arm historical calibration is not even reproducible at its own grain.
  5. Callee is an authored lexeme, not a resolved identity — a "ShellCommand" string literal and a ShellCommand constructor are the same atom.

gunbc.spark_managed_access_apply is the worked proof that these are capability gaps rather than API preferences. Its four script builders reach a host through four hops, one per axis: a list literal in a named argument, a match-arm binder, interprocedural flow into a service operation, and the transport stdin channel. No hop is visible. A detector over this surface finds them only by being told their names — the exact defect the cut exists to end.

What is delivered besides the blocker

The derived shell-interpretation seed, which is needed under whichever remedy lands, plus three findings I did not expect:

  • Of 262 transport shell declarations, nearly all are direct execs whose transport is merely named shell (find, test, mv, systemctl, jq). 17 sites carry a shell-interpreter argv literal — and 10 of those sit in product fn bodies as ArgvCommand constructions, one with a computed program, not in extdeps declarations. So axis 1 alone would find 6 and miss the majority; the declaration seed and real bodies are both required.
  • -- is overloaded three ways across 46 argv literals: ssh re-root, systemd-run re-root, cargo run -- argument pass-through, git ls-files -- end-of-options. No syntactic rule separates them; which reading holds is a per-program CLI contract, i.e. an extdeps citation duty (DESIGN §3). This is the one place in the seed where a new modeled authority is owed rather than a derivation, and it is StaticShellRouteUnknown until one exists.
  • The channel is selected by the flag, not the argv. -c puts the program in the next argv element; -s puts it on stdin. A census reading only argv misses every -s route, which includes shell.Exec.Run — the tree's main script route.

And one specimen found on the way, worth a row independent of this census: v2.compiler.source_authority canonical_dag_source_parse_print_law has zero callers anywhere in the tree — its name occurs exactly once, at its own definition. It is a parse/print law that nothing executes: DESIGN §5 specification-without-execution, in the compiler's own source authority. That is why "the full-fidelity parse route exists in .dag" is not evidence that it works on the real corpus.

Open question handed back

The full-fidelity route does exist unused (parse_dag_source_ast → ParseTree = Node, with Filesystem.List/Read already callable from .dag). Its viability is a measurement, not an inference, and the cost signal is adverse: ingested_fixture_arrows runs the same tokenize→parse→normalize→resolve path on a three-line module and is enrolled in the long lane for exceeding the 5s fast-lane budget. The census subject is 3,733 files and 761,844 lines. A probe against a real corpus file is running; the result decides whether the increment is "walk, read, parse, fold" or a genuine substrate cut, and per the dispatching lane the projection lands as its own PR with 0A rebased onto it either way.

Non-goals honored

No shell migration, no bridge deletion, no route disposition table (0B), no srv* decision, no stage0 mirror edit, no grep-generated population, no change to push_shell_argv_tokens (#8549 already closed that ambiguity; it survives as a negative control).


Measurement result (added after the probe): outcome 3, plus a correctness failure

The full-fidelity route was measured on both axes rather than assumed. Accepts-or-refuses was the wrong frame — there are three outcomes, and it lands on the third.

Correctness. With a positive control in the same run, the tree's own 3-line fixture ACCEPTS and a real 28-line corpus file ACCEPTS, so the harness is sound. On a random 10-file corpus sample: 6 accepted, 4 refused — and all four refusals carry the same reason as the earlier dag/extdeps/shell/exec.dag refusal, parse_grammar_choice_overlap_residue. The grouping is the finding, not the rate: one grammar deficiency with many victims. The route is a single repair away from a much larger accepted population, and equally no census can run on it until that repair lands — the merge bar is zero production parse refusals and the refusal set contains the census's own seed file.

Affordability, measured as a slope over a random 40-file sample rather than extrapolated:

files source bytes wall exit
1 372 63.1 s returned
10 5,332 67.5 s returned 6A/4R
40 388,527 187.3 s EXIT=137, OOM-killed after 34 files
  • Time is not the wall: ≈0.49 s marginal per file against ≈62.6 s fixed world-acquisition overhead — and that fixed cost is itself the shape under discussion.
  • Memory is, and not because of big files: the kill came at file ~34 with only 94 KB of source consumed, on ~7.8 KB files, with the 212 KB outlier sorted last and never reached. The fold accumulated only a short string, so this is not the probe's accumulator. Whether it is parse-tree retention or interpreter heap growth is not established by this probe and is not claimed — what is established is that the process cannot hold 34 small files, against a subject of 3,733 files and 31.3 MiB.

Recommendation. DESIGN §6 already rejected this shape in #8140 — "the unit of computation was the world, the unit of fact was one module's authorship" — and its declared next-rung trigger is one module's facts from one module's source, checked at ingestion where the module is already parsed. So the increment is the ingestion-side projection, not a census-side fold: a ServiceItem accessor (axis 1); a non-lossy body projection with labelled argument edges, arm/occurrence identity and resolved callee identity (axes 2–5); and a corpus-grain denominator that is an enumerated file set rather than an import closure (axis 6). 0A rebases onto it.

parse_grammar_choice_overlap_residue is worth filing separately either way: a named grammar deficiency refusing a large fraction of authored source in the compiler's own parser, invisible today because the only path that would surface it has no callers.


Two specs filed (no implementation started)

Per the dispatching lane: the increment lands partly in src/v1/stage0/src/coproduct_reflection.rs, which is the frozen v1 seed. Its admission test is purpose — does the change serve the v2 self-host program — and this increment serves a shell-migration census, a different program. So the call is the operator's and no seed work starts on lane authority.

docs/plans/parsed-body-projection-increment-spec.md — written for an admission decider, not an implementer:

  • Leads with the empty-observation narrow rather than the coverage percentage, because that is what makes the substrate unusable rather than merely partial: a file that was never loaded reads identically to a file that carries no shell route, and no per-file ParseRefused row exists to count the difference. A census on it produces a clean, reviewable, confident population that is silently wrong.
  • Each axis carries its own evidence grade — axes 2–3 execution-proven (the discriminating fixture pair), axis 6 execution-measured, axis 1 structural and independently verified, axes 4–5 read from the marshal.
  • The counting method is stated beside the number, since it will be questioned: 41,965 counts line-start fn/func/test fn and is the denominator matching the accessor's own FnItem || FuncItem filter — ItemKind has no separate test variant, so a test fn is an FnItem. 30,851 is the same count with the 11,114 test declarations removed. 1,698 visible is 4.0% or 5.5%; the conclusion is invariant.
  • Splits v1-seed from v2-.dag changes, and two properties matter for the call: axes 2–5 are an additive second accessor, not an edit to the existing marshal — whose lossiness is load-bearing for v2.lens.wiring_liveness and must not change — so the question is "may the seed grow an accessor", not "may its semantics change"; and axis 6 is probably already-sanctioned work pending P3+P5: infer whole-corpus scans to per-module maps #6239 rather than anything this increment requests.
  • States the admission question with both readings and no advocacy. If refused, the census stays blocked and honestly so — preferable to a census whose population is confident and silently wrong.

docs/plans/parse-grammar-choice-overlap-residue-finding.md — filed separately because it outlives this census. Five refusals across two source roots (four sampled at random, plus extdeps/shell/exec.dag found independently) all carry one reason. No corpus rate is claimed from ten files; the shared cause is the finding. It is invisible because the only path that would surface it has no callers — the unexecuted law and the unmeasured deficiency are the same fact seen twice.

🤖 Generated with Claude Code

https://claude.ai/code/session_01QFxNVPeTcYeCjP7rQyLtZu

gunbc-ci-auto-heal and others added 3 commits August 20, 2026 13:39
… the derived shell seed

The brief's stop condition fires. No detector was built and no text-scanning census was
substituted; this commit files the finding the brief asks for in that case.

The whole fact surface .dag can read is three accessors in coproduct_reflection.rs — types,
fns/funcs, data inits. ItemKind::ServiceItem has no accessor, so transport declarations are
invisible, and FnArrowDecl.output is a wiring-liveness skeleton whose statement fold DROPS a
let-bound RHS not referenced toward the return. Five axes with five distinct remedies, each with
its own evidence; the gunbc.spark_managed_access_apply chain is a worked proof that all five are
capability gaps rather than API preferences — a match-arm binder, a named argument label, a
service operation and a transport stdin channel, one per axis.

Also files the shell-interpretation seed derived from what interprets bytes as shell, which is
needed under whichever remedy lands. Findings that were not expected: of 17 shell-interpreter
argv literals, 10 sit in product fn bodies as ArgvCommand constructions (one with a computed
program), not in extdeps declarations, so a service-declaration projection alone would miss the
majority; and "--" is overloaded three ways across 46 sites (ssh re-root, systemd-run re-root,
cargo argument pass-through, git end-of-options), so the re-root reading is a per-program extdeps
citation duty and not something the census may infer from the token.

Records one specimen found on the way: v2.compiler.source_authority
canonical_dag_source_parse_print_law has zero callers anywhere in the tree — a parse/print law
that nothing executes, DESIGN §5 specification-without-execution in the compiler's own source
authority — which is why "the full-fidelity parse route exists in .dag" is not evidence that it
works on the real corpus.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QFxNVPeTcYeCjP7rQyLtZu
…easure the parse route

Axis 2 and axis 3 were argued from reading the Rust marshal; they are now proven by execution.
A fixture pair of identical shape folded through fn_arrow_decl_facts_live:
fixture_dead_let_shell (binds the call, does not use it) yields ZERO atoms — callee identity and
program literal both absent — while fixture_live_named_args yields
"fixture_sink2 | echo LIVEMARKER | echo ARGSMARKER". Same construct, opposite verdict, so the
loss is the projection's and not the probe's. The live arm also shows axis 3 directly: three
atoms in authored order with no labels, so nothing says which literal was program: and which was
args:.

Axis 6 is new and was measured, not read. The reflection registry is the ENTRY'S IMPORT CLOSURE,
not the corpus: a fold over fn_arrow_decl_facts_live under both production roots reported 1,698
fn/func declarations against 41,965 declared in the tree, about 4%. This is a declared frontier
rather than a discovery — corpus_dependency_view already refuses per-PR when
fn_arrow_decl_substrate_is_whole_tree is false ("blocked-on-#6239") — but it is fatal for a
census specifically, because files disappear through non-import with no per-file ParseRefused row
to count them. That is the empty-observation narrow: never-loaded is indistinguishable from
carries-no-route.

The full-fidelity parse route is measured against a positive control. The tree's own three-line
fixture ACCEPTS, a real 28-line corpus file ACCEPTS, and dag/extdeps/shell/exec.dag REJECTS with
reason=parse_grammar_choice_overlap_residue. So the route is real and its failure is a located
typed refusal, but the file it refuses declares shell.Exec.Run/RunArgv/Check — the census's most
load-bearing seed file.

Records that accepts-or-refuses was the wrong frame: there is a third outcome, accepts but is
unaffordable at corpus grain, and it is the one the prior cost signal makes likely. Affordability
is being measured as a slope over a random 40-file sample rather than extrapolated from the
fixture, since fixed overhead and per-byte cost are different curves. If it lands there, DESIGN §6
already rejected this shape once in #8140 — "the unit of computation was the world, the unit of
fact was one module's authorship" — and its declared next-rung trigger is exactly axis 1's
remedy: one module's facts from one module's source, checked at ingestion where the module is
parsed anyway.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QFxNVPeTcYeCjP7rQyLtZu
…come 3, and it fails correctness too

Accepts-or-refuses was the wrong frame; there are three outcomes and the measurement lands on the
third, with a correctness failure alongside it.

CORRECTNESS. On a random 10-file corpus sample the route accepted 6 and refused 4, and ALL FOUR
refusals carry the same reason as the earlier dag/extdeps/shell/exec.dag refusal:
parse_grammar_choice_overlap_residue. The grouping is the finding, not the rate — one grammar
deficiency with many victims rather than scattered file-specific problems. So the route is a
single repair away from a much larger accepted population, and no census can run on it until that
repair lands, because the merge bar is zero production parse refusals and the refusal set contains
the census's own seed file.

AFFORDABILITY, measured as a slope rather than extrapolated. 1 file / 372 B / 63.1 s; 10 files /
5,332 B / 67.5 s; 40 files / 388,527 B / EXIT=137, OOM-killed after 34 files. Marginal cost is
about 0.49 s per file against about 62.6 s of fixed world-acquisition overhead, so TIME IS NOT THE
WALL. Memory is, and it is not about big files: the kill came at file ~34 with only 94 KB of
source consumed, on ~7.8 KB files, with the 212 KB outlier sorted last and never reached. The
fold accumulated only a short result string, so the retention is not the probe's accumulator.
Whether it is parse-tree retention or interpreter heap growth is NOT established here and is not
claimed; what is established is that the process cannot hold 34 small files against a subject of
3,733 files and 31.3 MiB.

Both facts point the same way, and DESIGN §6 already rejected this shape in #8140 — "the unit of
computation was the world, the unit of fact was one module's authorship". Recommends the
ingestion-side projection over a census-side fold, enumerated as the six axes' remedies, with 0A
rebasing onto it. Notes parse_grammar_choice_overlap_residue as worth filing on its own: a named
grammar deficiency refusing a large fraction of authored source in the compiler's own parser,
invisible today because the only path that would surface it has no callers.

Two instrument faults on the way, same root and worth naming: a grep filter discarded every line
of a run, and a missing `bc` silently emptied the timing field while the surrounding output looked
healthy. Both failed toward absence, not toward a wrong number.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QFxNVPeTcYeCjP7rQyLtZu
@gunbai-bot
gunbai-bot Bot marked this pull request as ready for review August 20, 2026 14:17
Two specs, no implementation. The increment lands partly in the frozen v1 seed, whose admission
test is purpose — does the change serve the v2 self-host program — and this increment serves a
shell-migration census, so the call is the operator's and no seed work starts on lane authority.

The increment spec is written for someone deciding admission rather than for an implementer. It
leads with the empty-observation narrow rather than the coverage percentage, because that is the
part that makes the substrate unusable rather than merely partial: a file that was never loaded
reads identically to a file carrying no shell route, and no per-file ParseRefused row exists to
count the difference, so a census on it yields a clean confident population that is silently
wrong.

Each of the six axes carries its own evidence grade rather than being presented as uniformly
established — axes 2 and 3 execution-proven by the discriminating fixture pair, axis 6
execution-measured, axis 1 structural and independently verified, axes 4 and 5 read from the
marshal. The counting method is stated beside the axis-6 denominator because it will be
questioned: 41,965 counts line-start fn/func/test fn and matches the accessor's own filter, since
ItemKind has no separate test variant and a test fn IS an FnItem; 30,851 is the same count with
the 11,114 test declarations removed. 1,698 visible is 4.0% or 5.5% and the conclusion is
invariant.

Two properties are separated for the admission call: axes 2-5 are an ADDITIVE second accessor
rather than an edit to the existing marshal, whose lossiness is load-bearing for
v2.lens.wiring_liveness and must not change; and axis 6 is probably already-sanctioned work
pending #6239 rather than anything this increment requests. The admission question is recorded
with both readings and no advocacy.

The grammar finding is filed separately because it outlives the census. Five refusals across two
source roots, four sampled at random plus the independently-found extdeps/shell/exec.dag, all
carrying one reason: parse_grammar_choice_overlap_residue. No corpus rate is claimed from ten
files; the shared cause is the finding. It is invisible because the only path that would surface
it, canonical_dag_source_parse_print_law, has no callers — the unexecuted law and the unmeasured
deficiency are the same fact seen twice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QFxNVPeTcYeCjP7rQyLtZu
@briansrls
briansrls merged commit 6e50b24 into main Aug 20, 2026
1 check passed
@briansrls
briansrls deleted the session/sunny-pike-156 branch August 20, 2026 16:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant