Repository navigation
Frontier survey did not complete: interrupted-run receipt (0 receipts) + roster provenance rung-inflation finding - #7901
Conversation
…finding The authoritative 27/27 exact-head survey DID NOT RUN. This lands the executed refusal, not a measurement, and carries no frontier counts. frontier_probe_emit_from_ingest is OOM-killed before emitting a receipt for any module: five invocations, five kills, zero receipts. One-shot over all 27 (killed during module 1 of 27), plus 03_normalize, 03_resolve and 01_tokenize individually at --source-root src/v2 + dag, plus 03_normalize at src/v2 alone. Host cgroup memory.max = 33578549248 (31.3 GiB). All five die at the same place: after the per-module probe begins, inside the interpreted evaluation, with the ingest closure already fully read. Instrument control, because exit 137 alone proves nothing: cgroup oom_kill incremented 1 through 5, one per killed run, in lockstep. A 15-second RSS sampler saw only ~2.6 GB because the climb happens between samples; that figure is an artifact, not the working set. Two hypotheses raised and REFUTED by execution before being reported as causes: (a) that the 10 never-measured modules are the memory-bound ones -- refuted by 01_tokenize, one of the 17 that DID measure at 9f978aa; (b) that the widened src/v2 + dag root set is the cause -- refuted by 03_normalize OOMing at src/v2 alone. Cause remains open. Regression vs host capacity is UNRESOLVED and not resolvable here: three builds of the probe at old head 9f978aa were all OOM-killed (jobs=15, jobs=1, debug). The single-job attempt is decisive -- one rustc process still exceeded the cap, so this host cannot cold-build v1-compiler at all, and the current-head build succeeded only via sccache. Reported, not chased, per operator direction; a probe trimmed to fit a box measures a different question. Consequence: 27/27 is unreachable by BOTH documented paths, since a single module in a fresh process exhausts the cap alone. Fanning the 27 across workers does not address a per-module failure. Second, independent finding (DESIGN 4b rung inflation in the roster carrier): knowledge_attributed_blocker_class_seed_retained_row has ZERO call sites repo-wide, while all 10 never-measured roster modules are authored through execution_measured_seed_retained_row, whose name asserts execution measurement. That constructor takes blocker, stage and reason as ordinary parameters, so it cannot distinguish measured from attributed -- the NAME asserts a provenance the TYPE does not carry. No roster row is edited here; frontier.dag is load-bearing and the construction fix (a provenance field) is flagged for routing. Recipe corrections found by execution: ctrl-build defaults to remote and BuildBuddy refuses a detached worktree, so the mandated detached worktree and the default build path are mutually exclusive and ctrl-build --local is required. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… withdrawn
An independent same-cap control cold-built v1-compiler at a 3.26 GiB peak
with zero kills, refuting this receipt's earlier host-capacity claim. Re-reading
my own counters refutes the other one.
memory.max 33578549248 31.3 GiB
memory.peak 9136902144 8.5 GiB high-water across EVERY run
memory.events max 0 oom 0 oom_kill 10
max 0 means allocation NEVER reached memory.max; oom 0 means the CGROUP oom
killer was never invoked; in cgroup v2 oom_kill also counts tasks reaped by a
killer above this cgroup. The cgroup never exceeded 27% of its cap, so the kill
decisions were made outside its limit entirely.
WITHDRAWN (marked withdrawn in place, not deleted):
- "the probe requires more than 31 GiB" -> memory.peak is 8.5 GiB
- "this host cannot cold-build v1-compiler" -> refuted by the control
Root error, named so it does not recur: oom_kill incrementing in lockstep is a
sound control for "was this a kill" and worthless for "what hit the limit".
memory.events max/oom and memory.peak answer that, and were available throughout.
Not a leak either: memory.current was 391MB before the jobs=1 build with no
cargo/rustc running, and memory.peak never moved. A retry under low host load
survived to 3.5 minutes at ~2.0 GB, past every earlier run, then was killed
anyway without raising memory.peak. Cause is not attributable from inside the
container (/proc/self/cgroup is namespaced to 0::/); reported as an operator
finding rather than guessed at.
Consequently the lane is INTERRUPTED AND RETRYABLE, not blocked, and probe
memory cost is UNMEASURED rather than high. Doc and carrier note renamed off
"memory_refusal", which asserted the conclusion the evidence killed.
Both recipe-defect predictions also refuted by execution:
- no module collision: duplicate module in a third root does not refuse,
tested both inside and outside target/
- target/-resident roots ARE indexed: an invalid .dag there fails the compile
loudly (parse error naming the file, exit 101), so the emitted manifest
will be read
Only the ctrl-build --local defect survives. The vacuous-green property is
recorded separately: partition-totality passes on an empty manifest, so it can
never evidence that a survey ran.
Verification: all three edited .dag modules compile with zero own-file
diagnostics. Incidental and unrelated: 4 hard diagnostics in
dag/extdeps/systems/nvidia.dag (SPARK-0, byte offsets, DESIGN 4c body-grain
annotation refusal), present tree-wide and untouched by this branch.
Rung-inflation finding is unaffected and unchanged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Answers the question this lane could not resolve earlier: the ConstructorReferenceAdmission de-fork is #7902 (defork-constructor-admission). #7859 CREATED the fork it removes, so #7859 is not the de-fork and the gate must never be inferred from main moving. #7902 changes src/v1/04_infer.dag, so any frontier measurement taken before it merges is stale by construction. Records the operator-ordered sequence on the closeout carrier: interrupted-run receipt, then the FrontierMeasurementProvenance coproduct (so the survey does not rewrite roster rows through the provenance-conflating constructor this branch documents), then #7902, then the authoritative survey. Records the retry environment ruling: the survey must run in a fresh job or runner whose real kill boundary is observable, not in a session envelope reporting cgroup 0::/ that an ancestor can kill invisibly. The 27 modules shard, every receipt carries the same exact subject, and a killed shard is RERUN -- never hand-reconciled, since reconciling a shard that ran against a different subject is what turns a partial survey into a fabricated one. Verified: zero own-file diagnostics. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review found that commit c2edc86 corrected the survey carrier but left the LOCAL RECIPE text carrying the pre-correction reading -- a second, contradictory authority for the same facts, which is the DESIGN 3 single-authority violation this PR otherwise documents. Both findings verified against current code and both are real. gunbc.ci_layer_roots compiler_frontier_per_module_probe_exclusion_note: replaced "BLOCKED BY EXECUTED REFUSAL / OOM-killed / both paths unreachable / a single module exhausts the cap alone" with what the counters support -- runs interrupted, cause OPEN and not a demonstrated capacity limit (memory.events max 0, oom 0; memory.peak 8.5 GiB against a 31.3 GiB cap), probe memory cost UNMEASURED, "unreachable" NOT established, retry in a fresh job or runner with an observable kill boundary. The clause now points at frontier_probe_run_interruption_note rather than restating it, so there is one authority instead of two. gunbc.doc_graph_roots: dropped the "(memory refusal resolved)" parenthetical and "refusal receipt" wording from the bind dissolution; the trigger is the carrier's own -- the probe emits a receipt for at least one module. Swept the rest of the diff for the same staleness rather than fixing only what review cited, which surfaced two more the review did not mention: - the interruption note still opened "EXECUTED REFUSAL" and said frontier_probe_emit_from_ingest "is OOM-killed" - BOTH the note and the receipt said FIVE invocations; there were SIX. The low-host-load 01_tokenize retry -- the run that survived to 3.5 minutes at ~2.0 GB and is the strongest single piece of evidence against the capacity reading -- was described in prose but missing from the count and the table. An undercount in a receipt whose whole purpose is negative evidence is the same defect class review caught, so it is fixed here rather than left for a later reader to trip over. Verified: all three edited .dag modules compile with zero own-file diagnostics. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Both findings in You caught the exact defect this PR documents, committed by me while documenting it: 1 — 2 — Swept the rest of the diff for the same class rather than fixing only the two cited sites — which surfaced two more you did not flag:
An undercount in a receipt whose entire purpose is negative evidence is the same defect you caught, so it is fixed here rather than left for a later reader. Verified: all three edited |
…t jobs, four-PR gate
Three operator rulings supersede what this branch previously recorded.
NAMED SUCCESSOR, corrected shape. The provenance carrier is
FrontierEvidenceProvenance, not FrontierMeasurementProvenance, and
ExecutionMeasured takes a TYPED subject and a TYPED receipt rather than
free-form fields:
ExecutionMeasured { subject: FrontierSurveySubject, receipt: FrontierProbeReceipt }
KnowledgeAttributed { source: DeclarationRef, dissolve_on: DeclarationRef }
OracleDerived { oracle: DeclarationRef }
The row constructor must CONSUME the variant, so a row cannot claim
execution-measured provenance without carrying the receipt that proves it. That
turns this branch's rung-inflation finding from a documented defect into an
unwritable one -- a constructor named execution_measured_* that accepts ordinary
caller-provided blocker fields is a name-level claim, not evidence.
EXECUTION MODEL: 27 INDEPENDENT JOBS, NOT ONE PROCESS. One pinned subject, 27
independent module jobs, one receipt artifact per module, one exact-subject
aggregator. This is designed directly out of the failure this receipt records:
the costly property of the six runs was not that they died but that each death
cost EVERYTHING -- the one-shot lost all 27 modules when module 1 was killed. A
killed job now yields one interrupted row, rerun independently, and twenty-six
completed rows survive the twenty-seventh being SIGKILLed. A monolithic process
must not be retried in the same opaque envelope: the ancestor killer is still
unobservable from inside, so the answer is a visible kill boundary, not a bigger
cap.
GATE IS FOUR MERGES, NOT ONE: #7908 (whole-tree compile zero), #7902 (de-fork,
now READY at 227d9cb with its eight-control admission envelope green and the
arity ruling applied), #7925 (all-seven stage0 partition crate compile gate),
then this receipt. No authoritative survey runs on a head carrying compiler
diagnostics or a duplicated admission authority.
Also records the required output shape, whose last line is load-bearing: blocker
groups by DETAILED cause, not high-level reason strings. The 14 modules sharing
parse_grammar_choice_overlap_residue must have their detailed causes separated,
because the next phase ranks by largest DETAILED group on the premise that one
root governs many.
The receipt gets a pointer to the execution model rather than a copy of it; the
carrier stays the single authority.
Verified: zero own-file diagnostics.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…e-on with the doc bind Both findings in review 49609 verified against the tree and confirmed real. GATE ENUMERATION. The closeout note said "FOUR merges gate the authoritative survey" and "the gate opens after all four", then added a fifth prerequisite -- FrontierEvidenceProvenance -- in a later sentence of the same paragraph. Read as written, the four-merge list authorizes a survey run before provenance is structural, which would rewrite roster rows through the provenance-conflating constructor: precisely the rung inflation this receipt files. The note now says FIVE prerequisites and separates the two gates it was conflating -- the MERGE gate opens after the four merges, the SURVEY gate only after the fifth -- and states plainly that reading the four merges as the whole gate authorizes the defect. The duplicate provenance sentence downstream is deleted rather than left beside the corrected one, since two statements of one gate is how the disagreement arose in the first place. DISSOLVE-ON. frontier_probe_run_interruption_note dissolved on ONE condition (a probe receipt at a pinned subject) while the HandAuthoredDocBind for the same artifact requires TWO (receipt AND structural roster provenance). The bind is correct and the note was wrong: that bind carries frontier_roster_provenance_constructor_inflation_note as an additional work, so the weaker trigger would have deleted the receipt while the provenance finding it carries was still live -- a scaffold dissolving before its own trigger fires. The note now carries both conditions and names the bind as the single authority for when the doc deletes, so the two cannot drift apart again silently. Checked the md for a third copy: it states no dissolution, so there are exactly two authorities and they now agree. Verified: zero own-file diagnostics. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ight controls green Operator pass found two status claims in the closeout note that had rotted since they were written. Neither touches the interruption receipt or the roster provenance rung-inflation finding, which are the value of this PR and are unchanged. HEAD AND STATE. The note named #7902 as READY (not draft) at head 227d9cb. It is now DRAFT by operator ruling at head 42b5d00. ENVELOPE. The note claimed all eight admission controls green by execution. They are not, and the gap is a real defect rather than an unwritten test: writing control 2 exposed a fail-open where an unlisted caller can reference or invoke a ZERO-ARITY sealed constructor with no admission diagnostic at all, in both the bare-reference and the explicit-call spelling, proven by execution on two fixtures identical but for constructor arity. #7902 carries it as an inverted known-red, and the operator has ruled it is closed inside #7902 rather than merged as a known-red. Accurate statement is seven of eight green with the eighth quarantined pending that fix, so the gate does not open on a claim of eight. Recorded with the consequence for this lane attached: a survey taken against a head where a zero-arity sealed constructor is reachable without admission measures the wrong tree. Dropped the trailing "arity ruling applied" clause rather than carrying it forward unverified. That is the point of this correction - a status claim restated in a second place is a claim that rots silently, which is the same defect class this PR exists to record, so the note now asserts only the status the gate actually depends on. Gate structure is unchanged and stays separated as landed last round: four merges gate the merge, the FrontierEvidenceProvenance coproduct is the fifth and gates the survey. Verified: zero own-file diagnostics; no stale 227d9cb or eight-control reference survives anywhere in the tree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Operator ruling, superseding the two status patches it follows. The defect was not that the sentences were wrong; it was that the note was acting as a MUTABLE BOARD, so every merge made it wrong again and the fix arrived as another hand-patch. Two rounds of that is the signal. REMOVED from the closeout note: the four-PR gate list, the de-fork's head SHA, its draft state, and the seven-of-eight control tally. No PR number, no SHA, no lifecycle state, no tally survives anywhere in the file. REPLACED WITH conditions on the tree, which stay true as heads move and changes land. The authoritative survey may start only when all four hold: whole-tree hard diagnostics equals zero; ConstructorReferenceAdmission has one authority and no known-red control; all seven partition crates are continuously checked; FrontierEvidenceProvenance is construction-enforced. Each is decidable against whatever head is current, and for which change carries which condition the reader is sent to the board rather than to a second copy of it here. Two conditions are spelled out because they are easy to under-read. Condition 2 is not discharged by a de-fork landing while a control stays inverted -- one authority AND no known-red are both required, since a partially admitted constructor authority means the survey measures a tree in which a sealed constructor is reachable without admission. Condition 4 is what stops the survey writing all 27 roster rows through the conflating constructor. Why the reasoning is kept rather than just the list: a receipt that restates mutable status is a claim asserted in a place that cannot check it, which is the same defect class frontier_roster_provenance_constructor_inflation_note files one layer down. Recording that connection is what stops the next author reintroducing a status line here. Also sharpened the successor so it cannot read as landed implementation: the contract now says condition 4 is a construction wall to be BUILT and that this receipt does not build it -- this receipt RECORDS the provenance defect, the separate FrontierEvidenceProvenance change FIXES it. The PR body carries the same correction, retitled proposed shape NOT landed here, with the closing line changed from asserting the finding is already unwritable to stating that the defect stands recorded and unfixed until that change lands. Substantive content is untouched: zero receipts means no frontier measurement, the run was interrupted rather than proven memory-bound, the visible cgroup never reached its own limit, the survey needs independent module jobs, and the existing constructor inflates knowledge-attributed rows to execution-measured. Verified by execution: zero own-file diagnostics. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…updates Resolve ci.yml conflict by combining C1's all-seven partition cargo check gate (timeout 100) with main's heal-job generated-artifact exclusions. Co-authored-by: Cursor <cursoragent@cursor.com>
What this is, and what it is not
The authoritative exact-head 27/27 frontier survey did not run. This PR lands the executed receipt of an interrupted run plus one substantive finding. It contains no frontier measurement and no blocker-class counts, and it edits no roster row.
It also contains no probe memory requirement — see the corrections below.
The runs (6 probe invocations, 0 receipts)
Every probe invocation was SIGKILLed before emitting a receipt for any module: a one-shot over all 27 (died during module 1),
03_normalize,03_resolveand01_tokenizeindividually atsrc/v2+dag,03_normalizeatsrc/v2alone, and a later retry of01_tokenizeunder low host load. All die at the same place — after the per-module probe begins, inside the interpreted evaluation, with the ingest closure already fully read.Corrections: four hypotheses raised and refuted by execution
This PR's value is mostly in what got killed by controls rather than reported.
01_tokenize, one of the measured 17, fails toosrc/v2+dagroot set is the cause03_normalizefails withsrc/v2alonememory.peakis 8.5 GiBv1-compilerThe counter misread, named so it does not recur
max 0means allocation never reachedmemory.max.oom 0means the cgroup OOM killer was never invoked. In cgroup v2oom_killalso counts tasks reaped by a killer above this cgroup. The cgroup never exceeded 27% of its cap.oom_killincrementing in lockstep with the kills is a sound control for "was this a kill" and a worthless one for "what hit the limit". An earlier revision built a three-tier capacity narrative on exactly that inference.memory.events max/oomandmemory.peakanswer the second question and were available the whole time.Not a leak either:
memory.currentwas 391 MB before thejobs=1build with no cargo/rustc running, andmemory.peaknever moved. A retry under low host load survived to 3.5 min at ~2.0 GB — past every earlier run — then was killed anyway without raisingmemory.peak. A 2 GB process SIGKILLed with 80 GB free on the host, in a cgroup at 27% of cap, is not explicable by this cgroup, host exhaustion, or a session leak./proc/self/cgroupis namespaced to0::/, so the actual limiter is not observable from inside. Reported as an operator finding rather than guessed at.Consequently the lane is INTERRUPTED AND RETRYABLE, not blocked, and probe memory cost is UNMEASURED, not high. The receipt and carrier note were renamed off
memory_refusal— that name asserted the conclusion the evidence killed.Withdrawn claims are marked withdrawn in place rather than deleted, so the next reader sees the correction and the reasoning that produced it.
Finding: rung inflation in the roster carrier
Independent of everything above — a static fact about call sites and signatures, unaffected by any memory observation.
v2.compiler.self_host.frontierdeclaresknowledge_attributed_blocker_class_seed_retained_row— the constructor that honestly marks a row knowledge-attributed — and it has zero call sites repo-wide; the only occurrence is its own definition.Meanwhile all 10 roster modules never execution-measured at any head are authored through
execution_measured_seed_retained_row, whose name asserts execution measurement (8 directly;03_normalizeand03_body_producerviaquarantine_seed_retained_row_from_oracle, which delegates to it). The 10 are derived by joiningcompiler_frontier_sweep_orderagainst the interim TSV's module column, not counted by hand.The mechanism matters more than the count: that constructor takes
measured_blocker,located_stageandlocated_reasonas ordinary parameters, so it cannot distinguish measured from attributed — the name asserts a provenance the type does not carry, and no discipline at the call site can make the claim checkable. This is DESIGN §4b rung inflation, and it is why a carrier readingexecution_measuredis not evidence of measurement.The §5 construction fix is a provenance field making an attributed row unrepresentable as measured. No roster row is edited here —
frontier.dagis load-bearing, so this is flagged for routing, not landed.Recipe defects: one confirmed, two predicted and refuted
Confirmed:
ctrl-builddefaults to remote and BuildBuddy refuses a detached worktree (unexpected branch state * (no branch)). The recipe mandates a detached worktree, so the mandated procedure and the default build path are mutually exclusive;ctrl-build --localis required. Corrected in the recipe note.Refuted — no module collision: a duplicate manifest module in a third source root does not refuse, tested both inside and outside
target/.Refuted —
target/roots are indexed: an invalid.dagplaced in a third root undertarget/fails the compile loudly (parse error naming the file, exit 101), so the recipe's emitted manifest will be read.Recorded separately because it is real and independent:
compiler_frontier_wave2_blocker_partition_totality_holdspasses on an empty manifest, so it can never be read as evidence that a survey ran.Verification
All three edited
.dagmodules compile with zero own-file diagnostics.Incidental and unrelated to this branch: every compile at this head exits 1 on 4 hard diagnostics in
dag/extdeps/systems/nvidia.dag(SPARK-02be88489b13) — body-grain source annotations, the DESIGN §4c refusal. Cited offsets are byte offsets in a 120-line file. They appear even for a 26-source closure that excludes the file, so they are tree-wide. This diff touches four files and none isnvidia.dag.Changes
docs/probes/frontier_probe_run_interrupted_2026-08-06.md— the receipt (new)src/v2/compiler/self_host/frontier_probe_survey.dag—frontier_probe_run_interruption_noteandfrontier_roster_provenance_constructor_inflation_note, so the authority lives on the carrier rather than only in a doc (DESIGN §6: no parallel-ledger docs)dag/gunbc/doc_graph_roots.dag—HandAuthoredDocBindwith dissolution trigger (an unbound probe doc passes PR CI and reds the falsifier)dag/gunbc/ci_layer_roots.dag— recipe corrected with the--localrequirement and pointed at the interruption noteStatus
Ready for review; operator ruled the negative evidence should merge, because it prevents the next worker from increasing memory, narrowing source roots, or rerunning the same unobservable experiment and calling the result an OOM. It does not wait for the eventual successful survey — a negative result is publishable on its own.
Admission contract for the authoritative survey
Stated as conditions on the tree, naming no PR number, no head SHA, no draft state and no control tally. A receipt that names those is a second mutable copy of the board: wrong again every time the board moves, and patched by hand each time. That is the same defect class this PR's own finding is about, one layer down — a claim asserted somewhere that cannot check it — so this PR does not commit it.
The authoritative survey may start only when all four hold:
ConstructorReferenceAdmissionhas one authority and no known-red controlFrontierEvidenceProvenanceis construction-enforcedEach is decidable against whatever head is current. For which change currently carries which condition, read the board — not this PR.
Two are easy to under-read. Condition 2 is not discharged by a de-fork landing while a control remains inverted: one authority and no known-red are both required, because a partially admitted constructor authority means a survey measures a tree in which a sealed constructor is reachable without admission. Condition 4 is what stops the survey writing all 27 roster rows through the conflating constructor.
Named successor to the rung-inflation finding — proposed shape, NOT landed here
Operator-authored. This PR does not implement it: no type is declared, no constructor changed, no roster row touched.
#7901records the defect; a separate small carrier change fixes it, and must land before the survey (condition 4 above):The load-bearing detail:
ExecutionMeasuredrequires a typed subject and a typed receipt, and the row constructor must consume the variant — so a row cannot claim execution-measured provenance without carrying the receipt that proves it. A constructor namedexecution_measured_*that accepts ordinary caller-provided blocker fields is a name-level claim, not evidence. That is what would make the finding in this PR unwritable rather than merely documented — once the separate change lands. Until then the defect stands recorded and unfixed, which is the honest state.Survey execution model: 27 independent jobs, not one process
One pinned subject · 27 independent module jobs · one receipt artifact per module · one final exact-subject aggregator.
This is designed directly out of the failure this receipt records. The costly property of the six runs was not that they died but that each death cost everything — the one-shot lost all 27 modules when module 1 was killed. Under the ruled model a killed job yields one interrupted row, rerun independently; twenty-six completed rows are not lost because the twenty-seventh receives SIGKILL.
A monolithic 27-module process must not be retried in the same opaque session envelope: the ancestor killer is still there and still unobservable from inside, so the response is independent jobs with a visible kill boundary, not a bigger cap. Every receipt carries the same exact subject, and a killed job is rerun, never hand-reconciled — reconciling a job that ran against a different subject is what turns a partial survey into a fabricated one.
Required output shape when it runs
The last line is load-bearing and is not a table of high-level reason strings. The 14 modules that shared
^parse_grammar_choice_overlap_residueat9f978aa8dfneed their detailed causes separated, because the next phase picks the lowest-cost module in the largest detailed group on the premise that one root governs many — a high-level grouping cannot rank that.