Skip to content

Floor memory demand: uncensored peak, per-component attribution, the cost of a cap, and a standing a peak cannot travel without - #11743

Merged
gunbai-bot[bot] merged 11 commits into
mainfrom
session/proud-fox-12
Sep 20, 2026
Merged

gunbai-bot[bot] merged 11 commits into
mainfrom
session/proud-fox-12

Conversation

@briansrls

@briansrls briansrls commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

The measurement is complete, the judgment it motivates has landed with its witnesses green by
execution, and the instrument that takes the reading is bound and verified by invocation.

The instrument

gunbc test //gunbc/instruments:floor-memory-qualification runs the required-floor lane as a
child in the supervisor's own cgroup, waits, then reads memory.peak and memory.events from
outside the dead process and hands the raw counters to gunbc.floor_memory_demand to judge.

Why the child shares the cgroup — measured, not chosen. A systemd-managed unit reaps its
cgroup on exit, so the peak is gone before it can be read (Result=oom-kill survives; the peak does
not). A self-created cgroup cannot be joined by an existing process across a delegation boundary
(EPERM). Running the child in the cgroup the supervisor already occupies avoids both: nothing is
moved, and nothing reaps the cgroup because the supervisor still lives in it.

Why not an in-process guard. Drop covers ordinary returns and unwinding panics and nothing
else — not panic=abort, process::abort, process::exit or SIGKILL, and Rust's allocation-error
handler normally aborts. Allocation failure and the OOM kill, the two terminations this exists to
report, are exactly the ones no in-process observer can report.

The judgment stays in the substrate. The host reads bytes and knows nothing about what they
mean; qualify_floor_memory_from_readings takes primitives and interprets them in .dag over
extdeps.linux.cgroup_v2_memory's own parsers. A host that built CgroupMemoryLimitValue itself
would be a second parser for a file that authority already owns.

Two refusals that invoking it exposed

Both are checked before the workload, so a 35-minute run is never spent producing an
unattributable figure — and, more to the point, never tempts anyone to publish one anyway.

  • MeasurementCgroupShared — memory.peak is a property of the cgroup, not a process. An
    ordinary login session scope was measured holding five processes, so a bare invocation would
    report a neighbour's allocation as the floor's demand.
  • PeakDominatedByPriorHistory — memory.peak is the cgroup's lifetime maximum and this
    kernel refuses to reset it (EPERM, measured). A post-run peak that did not rise above the
    pre-run baseline belongs to something that ran earlier. The reset is attempted and read back
    rather than trusted.

Verified by invocation, not by compiling

There is no join between the .dag TargetProducer and the host's narrower Rust enum of the
same name, so adding a .dag variant forces no host arm and the build stays green.
RequiredFloorProducer is the standing proof: one occurrence corpus-wide, its own declaration, no
binding, no arm, no consumer. (I had claimed the compiler would catch a half-done registration; it
does not. Deliberately left dangling rather than repurposed — what this produces is a memory
qualification, not a floor verdict.)

So registration was checked by running it:

All three arms exercised, same binary and tree, one variable:

arm conditions result exit
DemandObserved MemoryHigh=infinity peak=15746146304, events all zero 0
DemandBounded MemoryHigh=8G peak=8590430208, high events 20242 1
refused bare, in a shared session scope MeasurementCgroupShared, pids named 2

The bounded run reproduces the censoring signature exactly: its peak sits 495616 bytes above the
8589934592 throttle line — the reclaim clamp's overshoot, allocations crossing in a burst while
direct reclaim drags usage back. Peak-at-the-line is the shape every figure behind the old 16/15
row had, and the instrument now refuses to call it a demand.

Both runs exited 1 from the floor itself — main is currently red on 24 annotation errors in
src/v2/workflow/floor_naming_hygiene.dag, a file this branch does not touch — so these peaks are
short runs', not comparable floor figures. That is the instrument reporting correctly: an ordinary
non-zero exit does not withhold completion, so the reading's standing is decided by the counters
rather than by the workload's verdict.

The judgment it rests on

gunbc.floor_memory_demand — a peak may not travel without its censoring standing, derived from
the kernel's own counters rather than authored beside them. Three arms, and which three is the point:

arm meaning
DemandObserved nothing held it — the only arm citable as a demand
DemandBounded a lower bound, whether the process was killed or throttled; cause in the fields
DemandUnreadable the supervisor could not read — must never render like either

Killed and throttled share one arm because the consequence for a consumer is identical: the number
may not be used to size anything downward. A supervisor that cannot tell could not read from it
fit
reproduces one layer out the conflation this exists to remove.

Two axes kept apart. Termination (how the process ended, observed from outside) and censoring
(whether anything held it) are independent — Run B below exits cleanly having been throttled
throughout, and a killed run may never have been throttled at all.

Why the subject is the process and the reading is taken from outside. An in-process observer — a
Drop guard, an exit hook, a final log line — covers ordinary returns and unwinding panics and
nothing else. It does not run on panic=abort, process::abort, process::exit or SIGKILL, and
Rust's allocation-error handler normally aborts. So allocation failure and the OOM kill, the two
terminations a memory instrument most exists to report, are exactly the ones no in-process observer
can report. Measured 2026-09-19: a systemd-managed unit reaps its cgroup on exit (peak gone,
though Result=oom-kill survives), while a cgroup the supervisor created itself keeps memory.peak
and memory.events readable past a SIGKILL. A process also cannot be moved across the
delegation boundary (EPERM) — it must be spawned into the cgroup.

test.claim.floor_memory_demand_witness — six witnesses, all green by execution via
claim_batch, not typechecked-and-assumed. The pair is two real runs (below). Four more cover what
counters alone would miss: a peak pinned exactly to the line with zero events — the shape every
figure behind the old 16/15 row had, and the case that reads as a demand if you consult only
memory.events — plus a kill far below every limit, an unobserved termination, and an unreadable
cgroup.

gunbc.recurring_failure_mode.suppressed_precondition_failure_runs_the_workload_unconstrained —
filed from a near-miss I caused taking these measurements: a cgroup placement write failed EPERM
under a 2>/dev/null, so a probe ran unconstrained on a shared host. An uncapped run is
indistinguishable from a capped one in everything the workload itself emits, so the precondition must
be verified by readback, never inferred from the call's exit status. Checked against the two
nearest existing rows before minting: control_plane_acknowledgement_minted_as_effect excludes it in
its own words (it requires that every precondition held and no diagnostic fires), and
absorbing_fallback does not fit since nothing widened and no value was substituted. The repo's own
ctrl-build "forwarding env: (none)" warning is the same class on a different knob.

dag/gunbc/live_deploy/emit.dag — citation fix; see the bottom of this body.

The measurement these rest on

Two runs on srv1 at 56375ec44a (= origin/main), same binary, same tree, back to back, one
variable. Method note: the recipe's high=max is the cgroup readback spelling;
systemd-run -p MemoryHigh=max is rejected outright (Failed to parse MemoryHigh=max) and the
systemd spelling is infinity. A recipe transcribed literally either fails to launch or, if the
property is dropped, produces a CONSTRAINED run that gets reported as uncensored.

Run A — uncensored (MemoryMax=64G, MemoryHigh=infinity, MALLOC_ARENA_MAX=2):
peak 28962353152 B = 26.973 GiB, scope memory.events all zero, stall 0 faults/min.
Nothing held it, so it is a demand rather than a pin. That is above the current slot
(memory_max 26.000 GiB, memory_high 25.000 GiB) by 0.973 and 1.973 GiB.

Run B — 16 GiB throttle line: peak 17181028352 = 16.001 GiB, i.e. memory.high exactly —
the censored signature, reproduced deliberately under a known limit. 17466 high events,
623358 faults/min.

A 16 GiB cap costs +23.0% of wall, and the fold is the phase that hides it

Run A Run B
whole run 2072 s 2548 s +23.0%
phase cost phase cost
declarer-discovery +1.4% discovery-authority +35.2%
strict-preparation +6.7% site-projection +34.6%
gate-closure +7.2% whole-tree-graph-facts +77.8%
claim fold +8.1% published-mock-projection +76.8%

The phases that run once preparation has filled memory — where the live set sits farthest above
the line — pay 35–78%. A cap decision denominated in fold duration reads the one phase that
does not show the cost.
Sharper: my uncensored fold took 726508 ms, within 0.07% of the
727002 ms previously attributed to a throttled fold on another host. Across hosts, contention
dominates that number and the memory variable is not isolated even in principle.

What the sampler is still for

memory.peak is a kernel-maintained maximum, so the peak needed no sampling — the supervisor
design reads it once, post-mortem. That does not make the 2-second sampling unnecessary: the
phase attribution table above and the RSS-versus-charge comparison below both come from the paired
samples, and neither is obtainable from a single post-mortem read. Two separate products, one of
which the supervisor replaces and one of which it does not.

Attribution

segment RSS at end Δ
floor entry 6.58 GiB
gate-closure ×2 12.30 +5.72
rest of strict-preparation 24.75 +12.45
module-index → published-mock 24.75 +0.00
discovery-authority / site-projection 24.76 +0.01
claim fold (4184 claims) 26.70 +1.94
local-repo-wet lane 26.97 +0.27

strict-preparation is +18.17 GiB, 89% of floor-entry-to-peak growth. The peak lands in the
local-repo-wet lane
, a fourth component the preparation/fold/teardown split does not name.

pool_parse is a deliberate negative. pool-root-index-warm ran 456 ms with
rss_growth_bytes=0, and every warm sub-phase reports 0 or a few hundred KiB. The phase named
for it pays nothing because something upstream already forced the parse, and nothing in the
floor's output reports that forcing — there is no pool_parse accumulator and no [calibration]
line in this lane. No byte figure is derivable from this run and none is offered.

Scope-vs-charge: a level difference, not a counter difference

On ONE cgroup over 967 paired samples, memory.current − RSS was mean 25.3 MiB, max 105.1 MiB,
max relative gap 1.33%. The 2026-09-12 pair differ by 1.96 GiB (8.6%) — two counters on one
cgroup do not produce that. The ancestor walk does, same instant: scope peak 26.97 GiB with all
events zero; app.slice 197.14 GiB; user-1000.slice 249.52 GiB, carrying max 1854 and
oom_kill 2 that belong to other descendants. This identifies the only mechanism that produces
charge-above-scope; it does not retro-diagnose that pair, whose note records no level for
either figure.

Honest caveats

Both runs ended exit=1, phases_run=3 phases_failed=2 — the same two local-repo-wet self-host
witnesses, after the fold — so they are directly comparable to each other but neither is
FloorClean. Because those two refused early in the lane where the peak occurs, 26.973 GiB is a
lower bound on a clean run. srv1 was at load average 55–75 with 400+ users throughout, so no
absolute duration here is comparable to CI; only A against B.

What remains: the supervisor

floor_cgroup_envelope is documented as "Emitted at entry and again at exit so the peak and the
event counters bound the whole run."
There is no exit call — the only call sites are
"floor-entry" and the heartbeat's every-tenth-beat "beat-N". At CI's 60s cadence (unset in the
workflow; std/observation.dag independently records "the floor still beats once a minute") the
envelope fires every 600 seconds and never at exit. So the floor cannot report its own peak,
which is why an external sampler was needed here and is a sufficient explanation for how a note
could end up with two figures and no record of which level either named.

The censoring standing has landed (gunbc.floor_memory_demand, above), so a reading can no
longer travel without it. What remains is the thing that takes the reading: a supervisor that
creates the cgroup, spawns the floor into it, waits, then reads memory.peak, memory.events and
the wait status from outside the dead process. That is the arm immune to abort, non-unwinding panic,
process::exit and SIGKILL, and it is why the supervisor — not an in-process guard — is the
load-bearing evidence. An exit envelope inside the floor remains worth adding as a convenience that
supplies phase-correlated internal endpoints, but it can never be the reading, and it must be a
scope guard rather than a placed call: run_required_floor is 4643 lines with 48 return Err(
sites and 59 ? propagations against zero explicit return Ok(, so a hand-placed exit emission
would cover the success path and miss all 107 refusal paths — which are the runs the instrument
exists for, by that file's own comment.

Not in scope here: the slot values in gunbc.runner_slot_allocation. The peak exceeds the
current row, but re-sizing is a capacity decision with a CPU-width consequence and is being
handled in a separate lane.

🤖 Generated with Claude Code

The passage quoted "22.74 GiB uncensored" as what the 2026-09-12 demand ruling
sized the runner slot from. The ruling says the opposite in as many words: it
records TWO disagreeing readings, takes the LARGER, and calls sizing to the
smaller "the fail-open direction". So this row attributed to the ruling the exact
arm the ruling refused, and named the fail-open direction as what the fail-closed
decision was.

The figure is REMOVED rather than corrected. This passage's own argument is that
the floor's demand does not apply to this deployment at all, and that holds at any
value -- so nothing downstream of it changes. It is a citation defect, not a wrong
decision. Carrying a second copy of a number owned by
gunbc.runner_slot_allocation gunbc_runner_slot_memory_max_ruling_note is the
DESIGN section 6 defect (name the producer, never transcribe its output), which is
also why no newer figure is substituted here even though one now exists.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gunbai-bot gunbai-bot Bot changed the title Floor memory: an uncensored demand measurement and a per-component attribution Floor memory: uncensored demand, per-component attribution, and the cost of a cap (WIP — citation fix only so far) Sep 19, 2026
…ding

THE DEFECT THIS CLOSES. A peak read from a run that was HELD measures what the
process was ALLOWED, not what it needs. gunbc.runner_slot_allocation already
states that reading -- a peak equal to the throttle line is "the censored lower
bound ... the workload being held, not the workload being measured" -- and then,
on 2026-09-12, recorded two disagreeing figures for one run and said it did not
know which cgroup level each named. Nothing in the corpus made that ambiguity
unwritable, so the smaller figure propagated to gunbc.live_deploy.emit as "the
required floor's measured peak, uncensored".

gunbc.floor_memory_demand derives the standing from the kernel's own counters
instead of leaving it to prose. Three arms, and which three is the whole point:
DemandObserved (nothing held it -- the only arm citable as a demand),
DemandBounded (a lower bound, whether because the process was killed or because
it was throttled -- one arm because the consequence for a consumer is identical,
with the cause in the fields), and DemandUnreadable, which must never render like
either. A supervisor that cannot tell "could not read" from "it fit" reproduces
one layer out the conflation this exists to remove.

TWO AXES, KEPT APART. Termination (how the process ended, observed from outside)
and censoring (whether anything held it) are independent: a run can exit cleanly
having been throttled throughout, and a run can be killed having never been
throttled. Folding them into one enum loses exactly the pair an operator needs.

WHY THE SUBJECT IS THE PROCESS AND THE READING IS TAKEN FROM OUTSIDE. An
in-process observer -- a Drop guard, an exit hook, a final log line -- covers
ordinary returns and unwinding panics and nothing else. It does not run on
panic=abort, process::abort, process::exit or SIGKILL, and Rust's
allocation-error handler normally ABORTS. So allocation failure and the OOM kill,
the two terminations a memory instrument most exists to report, are precisely the
ones no in-process observer can report. The kernel maintains memory.peak and
memory.events regardless; a supervisor that outlived the child can still read
them. Measured 2026-09-19: a systemd-managed unit REAPS its cgroup on exit (the
peak is gone, though Result=oom-kill survives), while a cgroup the supervisor
created itself keeps both readable past a SIGKILL.

THE DISCRIMINATING PAIR IS TWO REAL RUNS, NOT TWO FIXTURES. Both on srv1 at
56375ec, same binary, same tree, back to back, one variable -- the throttle
line. Run A uncensored: peak 28962353152, all events zero, qualifies as a demand.
Run B at a 16 GiB line: peak 17181028352 with 17466 high events, does not. Four
further witnesses cover the cases counters alone would miss: a peak pinned
exactly to the line with zero events (the shape every figure behind the old 16/15
row had), a kill far below every limit, an unobserved termination, and an
unreadable cgroup. All six green by execution via claim_batch.

ALSO: one recurring_failure_mode row, filed from a near-miss I caused taking
these measurements. A cgroup placement write failed EPERM under a 2>/dev/null and
the workload ran UNCONSTRAINED on a shared host -- and an uncapped run is
indistinguishable from a capped one in everything the workload itself emits. The
row's general form is that a precondition establishing the ENVIRONMENT must be
verified by readback, never inferred from the call's exit status; the repo's own
ctrl-build "forwarding env: (none)" warning is the same class on a different knob.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gunbai-bot gunbai-bot Bot changed the title Floor memory: uncensored demand, per-component attribution, and the cost of a cap (WIP — citation fix only so far) Floor memory demand: uncensored peak, per-component attribution, the cost of a cap, and a standing a peak cannot travel without Sep 19, 2026
Brian Searls and others added 3 commits September 20, 2026 03:37
…ost arm pending)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…group read

THE INSTRUMENT. `gunbc test //gunbc/instruments:floor-memory-qualification` runs the
required-floor lane as a CHILD IN THE SUPERVISOR'S OWN CGROUP, waits, then reads
memory.peak and memory.events from outside the dead process and hands the raw
counters to `gunbc.floor_memory_demand` to judge. The three terminations map onto
the three arms: DemandObserved is 0, DemandBounded is 1 (a lower bound, which may
not size anything downward), every refusal is 2.

WHY THE CHILD SHARES THE CGROUP. Measured, not chosen: a systemd-managed unit REAPS
its cgroup on exit so the peak is gone before it can be read, and a self-created
cgroup cannot be JOINED by an existing process across a delegation boundary (EPERM).
Running the child in the cgroup the supervisor is already in avoids both -- nothing
is moved, and nothing reaps the cgroup because the supervisor still lives in it.

THE JUDGMENT STAYS IN THE SUBSTRATE. The host reads bytes and knows nothing about
what they mean; `qualify_floor_memory_from_readings` takes primitives and does every
interpretation in .dag over `extdeps.linux.cgroup_v2_memory`'s own parsers. A host
that built CgroupMemoryLimitValue itself would be a second parser for a file that
authority already owns.

TWO REFUSALS ADDED AFTER INVOKING IT EXPOSED THEM, both checked BEFORE the workload
so a 35-minute run is never spent producing an unattributable figure:

  MeasurementCgroupShared -- memory.peak is a property of the CGROUP, not a process.
  An ordinary login session scope was measured holding FIVE processes, so a bare
  invocation would report a neighbour's allocation as the floor's demand. Verified by
  execution: invoked in a session scope it refuses with exit 2 and names the pids.

  PeakDominatedByPriorHistory -- memory.peak is the cgroup's LIFETIME maximum and this
  kernel REFUSES to reset it (EPERM, measured). A post-run peak that did not rise above
  the pre-run baseline belongs to something that ran earlier, so it is refused rather
  than attributed to this run. The reset is attempted and READ BACK rather than trusted.

WHAT THE COMPILER DOES NOT CATCH, recorded because I claimed otherwise and was wrong:
there is no join between the .dag TargetProducer and the host's narrower Rust enum of
the same name, so adding a .dag variant forces NO host arm and the build stays green.
`RequiredFloorProducer` is the standing proof -- one occurrence corpus-wide, its own
declaration, no binding, no arm, no consumer. Registration was therefore verified by
INVOKING the label, not by compiling: the target now appears in `gunbc test`'s
available-target list and its refusal arm runs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ledger-Repair-Judged: docs/design-rung-drops.md
Ledger-Rows-Repaired: docs/design-rung-drops.md namespace_wave_admission_wall_removed
Heal-Candidate-Run: 35487819180
@gunbai-bot
gunbai-bot Bot marked this pull request as ready for review September 20, 2026 04:33
Brian Searls and others added 3 commits September 20, 2026 05:29
… rows nothing read

FINDING 1 (DESIGN section 7) — ACCEPTED AND FIXED. A 280-line hand-Rust host
module plus its host declarations landed with no row in
gunbc.target_invocation_seed_growth. The row now carries all nine
floor_memory_supervisor items, the two subject accessors and
run_floor_memory_qualification, and current_boundary names the new file and the
mod line.

ITS TRIGGER IS STATED SEPARATELY AND IS HONESTLY WEAKER than the rest of the row,
because assimilating it to the existing sentence would promise a discharge that is
not in sight. The other subsets dissolve when the emitter reaches their modeled
modules. This one does not: every operation in it is a resource effect — read
/proc/self/cgroup, spawn a child, wait on it, read /sys/fs/cgroup after the child
is gone — and resource operations resolve only inside a workflow function realized
by the seed INTERPRETER, which an emitted binary is not running under. It
discharges when the substrate can express a supervised child process as modeled
effects with typed refusals, and not before. Moving the judgment further into .dag
does not discharge it, and neither do witnesses over the read.

FINDING 2 (DESIGN section 3c) — ACCEPTED, AND THE FIX IS THE REVIEW'S SECOND
OPTION FOR A REASON I HAD TO MEASURE. floor_memory_qualification_source_roots and
floor_memory_qualification_lane had their only occurrences at their own
definitions, in the same file whose new comment states that exact test. The review
offered two remedies: route them through a real consumer, or drop them and let the
Rust own the fact explicitly.

I BUILT THE FIRST ONE, AND IT CORRUPTS THE MEASUREMENT. Reading those rows means
resolving a corpus graph, and this supervisor shares its cgroup with the child BY
DESIGN — that sharing is what lets memory.peak survive the child's death. So the
resolve lands in the very counter the instrument reports. Same failing floor, same
tree, same binary, differing only in whether the subject was read from the model:

  Rust-owned subject   peak 15746146304  (14.66 GiB)
  Rust-owned subject   peak 15704227840  (14.63 GiB)
  read from the model  peak 22293544960  (20.76 GiB)   <- +6.1 GiB, 42% inflation
  Rust-owned again     peak 15816912896  (14.73 GiB)   <- restored

An instrument may not consult the authority from inside the cgroup it measures: the
act of reading perturbs the reading. Netting the supervisor's footprint back out
was not available either — that replaces a measured number with an adjusted one,
which is the habit gunbc.floor_memory_demand exists to refuse.

So the two rows are DELETED rather than left unconsumed, and the Rust says plainly
that it owns the subject and why, with the figures above and with the condition
that would restore the modeled form: any route that reads them OUTSIDE the measured
cgroup. That needs the cgroup lifecycle modeled, which is the same capability the
seed-growth trigger names.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 20, 2026

Copy link
Copy Markdown
Contributor

Both findings in review 68906 verified and fixed in 8245f6f. Thank you — the second one was right in a way I would not have found on my own, and chasing it produced a measurement worth having.

Finding 1 (§7) — accepted, fixed. gunbc.target_invocation_seed_growth now carries all nine floor_memory_supervisor items plus the host declarations, and current_boundary names the new file and the mod line.

Its trigger is stated separately and is honestly weaker than the rest of the row. The other subsets dissolve when the emitter reaches their modeled modules; this one does not. Every operation in it is a resource effect — read /proc/self/cgroup, spawn a child, wait on it, read /sys/fs/cgroup after the child is gone — and resource operations resolve only inside a workflow function realized by the seed interpreter, which an emitted binary is not running under. It discharges when the substrate can express a supervised child process as modeled effects with typed refusals, and not before. I have said explicitly that moving the judgment further into .dag does not discharge it, and neither do witnesses over the read, so the row cannot be retired by work that leaves the four effects standing.

Finding 2 (§3c) — accepted, and I took your second option because I measured the first one failing.

You offered: route the two facts through a real consumer, or drop them and let the Rust own the fact explicitly. I built the first. It corrupts the measurement.

Reading those rows means resolving a corpus graph, and this supervisor shares its cgroup with the child by design — that sharing is precisely what lets memory.peak survive the child's death. So the resolve lands in the very counter the instrument reports. Same failing floor, same tree, same binary, differing only in whether the subject was read from the model:

subject source peak
Rust-owned 15746146304 14.66 GiB
Rust-owned 15704227840 14.63 GiB
read from the model 22293544960 20.76 GiB — +6.1 GiB, 42% inflation
Rust-owned again 15816912896 14.73 GiB — restored

An instrument may not consult the authority from inside the cgroup it is measuring; the act of reading perturbs the reading. Netting the supervisor's footprint back out was not available either — that replaces a measured number with an adjusted one, which is the habit gunbc.floor_memory_demand exists to refuse.

So the two rows are deleted rather than left unconsumed, and the Rust now says plainly that it owns the subject, why, with those figures, and what would restore the modeled form: any route that reads them outside the measured cgroup. That needs the cgroup lifecycle modeled — the same capability the seed-growth trigger names.

You were also right that my own comment condemned the diff. I have replaced that comment's name-list with the rule and the check, since enumerating which siblings dangle goes stale the moment anyone adds an instrument.

On the refuted pre-scan row: agreed, and thank you for refuting it rather than echoing it.

One note on docs/design-rung-drops.md: your read was right. I re-derived it — the deleted POSTSCRIPT, 2026-09-20 has no occurrence in the dag/ authority or on main, so the regeneration is correct-direction. The file came back as an unmerged path with no conflict markers when I merged main (the refusing generated-artifact driver), and I regenerated it through tools.docs_projection_gate regen rather than hand-resolving it.

— sent from proud-fox-12

Brian Searls and others added 2 commits September 20, 2026 06:18
…e fail-open

FOUR FINDINGS, ALL VERIFIED AGAINST THE CODE BEFORE FIXING.

1. THE CONTAMINATION GUARD FAILED OPEN ON THE CASE IT WAS BUILT FOR, and this is
the serious one. `cgroup.procs` lists DIRECT members; `memory.peak` aggregates the
whole SUBTREE. `nearest_cgroup_with_peak` deliberately climbs to an ancestor, so
the guard could find no strangers in a cgroup whose descendants were running
anything at all. Measured on srv1 while confirming it: user-1000.slice has ZERO
direct processes — a direct-membership check finds nothing — against 242 child
cgroups, 336 processes beneath it, and memory.peak 392042180608. The instrument
would have reported a third of a terabyte of co-tenant allocation as the floor's
demand, as DemandObserved.

Fixed twice over: membership is now walked over the SUBTREE, and a cgroup with any
child at all is refused separately, because a descendant can be created after the
check and only a leaf is stable. EXERCISED, not asserted: in a delegated scope with
a deliberately created child cgroup the instrument refuses with
MeasurementCgroupHasChildren and exit 2.

2. THE REFUSAL VOCABULARY HAD FORKED IN BOTH DIRECTIONS. `DemandReadRefusalCause`
carried three arms nothing could construct (a check whose forbidden state is
unwritable is a decoration, section 4b) while the producer carried five real causes
the model had never heard of — and every refusal the instrument actually emits came
from the unmodeled set.

The repair is not to copy one list into the other: THE TWO POPULATIONS HAVE
DIFFERENT SUBJECTS. A supervisor's refusals happen BEFORE there is anything to
judge — no cgroup, no child, no attributable counter — and terminate the invocation
with no observation, never reaching the fold. What reaches the judgment is a
complete set of readings, so the only way IT can refuse is that a reading cannot be
interpreted. One arm, because there is exactly one such way. The Rust comment
claiming to mirror the model is corrected to say the opposite and why.

`TerminationUnobserved` is kept and its reachability stated plainly: exercised by
the witness, not by today's producer, because the fail-closed reading it encodes is
a property of the judgment rather than of one producer.

3. THE SEED CENSUS WAS ITEM-INCOMPLETE — five private helpers omitted, including
`nearest_cgroup_with_peak`, which carries the ancestor-climb decision behind finding
1. Privacy is not the grain this roster uses; it already enumerates private helpers
of the sibling module. Six added.

4. `peak_bytes: Int` — two reviewers split on this row (68906 refuted it as a
correct host-to-substrate primitive boundary, 68936 echoed it as unrefuted). It now
carries an explicit dissolve-on gate rather than an argument, stating why the
parameter is primitive at this one seam, that the scalar does not propagate, and
what would dissolve it: an interpreter argument surface that admits a constructed
ByteSize. Explicitly NOT dissolved by wrapping the literal one frame outward.

A new witness drives the refusal through the real entry point rather than
hand-constructing the variant, so the arm is evidence about the production route.
All seven witnesses green by execution.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… as unbound

THE FINDING IS CORRECT AND THE DEFECT WAS MINE. `peak_reached_limit` answered
`false` for `CgroupMemoryLimitUnparseable` -- "this line did not bind" -- so a
memory.max or memory.high body the upstream parser rejects, on a run that exited
cleanly with zero events, published DemandObserved: the one arm whose peak may be
cited as a DEMAND, for a run whose throttle line was UNKNOWN. A widen where DESIGN
section 5 requires a refusal, and the module's own comment conceded the premise it
then violated ("an unparseable one cannot be reasoned about at all").

FIXED IN TWO PLACES, deliberately.

The primary wall is a pre-check: an unparseable limit now refuses to
DemandUnreadable { CgroupValueUnparseable } before any other question is asked,
naming WHICH file and carrying the body the operator needs to see. Every question
downstream -- was it pinned, did it fit -- is a comparison AGAINST the limits, so a
limit that could not be read makes those unanswerable rather than negative.

Defence in depth is the second: the Unparseable arm of `peak_reached_limit` now
answers `true` rather than `false`. If the guard above it were ever bypassed, the
fail-closed reading is "treat it as pinned" -- DemandBounded, a lower bound nobody
may size down from -- and never DemandObserved.

THE RED IS PROVEN DISCRIMINATING, NOT ASSUMED. Three new claims drive unparseable
bodies through the real host entry point, one per limit file plus one asserting the
outcome is DemandUnreadable specifically (DemandBounded would also answer "not a
demand", so without that third claim the repair could have been a silent downgrade
for a reading that cannot be interpreted at all). I reverted the fix and ran them:
all three FAIL on the old code and PASS on the new. Ten witnesses green.

ALSO FROM THIS REVIEW: the `peak_bytes: Int` pre-scan row is now recorded as
REFUTED by the reviewer, on the dissolve-on gate added in the previous commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 20, 2026

Copy link
Copy Markdown
Contributor

Review 68949 verified and fixed in d848725. The blocking finding was right and the defect was mine.

peak_reached_limit answered false for CgroupMemoryLimitUnparseable — "this line did not bind" — so a memory.max/memory.high body the upstream parser rejects, on a run exiting cleanly with zero events, published DemandObserved: the one arm whose peak may be cited as a demand, for a run whose throttle line was unknown. A widen where §5 requires a refusal, and the module's own comment conceded the premise it then violated.

Fixed in two places, deliberately:

  • Primary wall — an unparseable limit now refuses to DemandUnreadable { CgroupValueUnparseable } before any other question is asked, naming which file and carrying the body. Every downstream question is a comparison against the limits, so an unreadable limit makes them unanswerable rather than negative.
  • Defence in depth — the Unparseable arm of peak_reached_limit now answers true, not false. If the guard above it were ever bypassed, the fail-closed reading is "treat it as pinned" → DemandBounded, never DemandObserved.

The red is proven discriminating, not assumed. You noted correctly that no claim drove an unparseable limit body. There are now three, through the real host entry point: one per limit file, plus one asserting the outcome is DemandUnreadable specifically — because DemandBounded would also answer "not a demand", so without it the repair could have been a silent downgrade for a reading that cannot be interpreted at all.

I reverted the fix and ran them:

FAIL an_unparseable_high_limit_is_not_a_demand_holds
FAIL an_unparseable_max_limit_is_not_a_demand_holds
FAIL an_unparseable_limit_lands_in_demand_unreadable_holds

and with the fix restored, all ten witnesses pass. So your failure scenario was reachable and is now walled.

On the peak_bytes: Int row — thank you for recording the refutation rather than echoing it. For the record of how it resolved: review 68906 refuted it, review 68936 echoed it as unrefuted-by-the-diff, and rather than argue I added the explicit dissolve-on gate the scan itself offers, naming what would dissolve it (an interpreter argument surface admitting a constructed ByteSize) and what would not (wrapping the literal one frame outward).

— sent from proud-fox-12

…(review 69097)

The 🟡 gate carried a reason, a non-propagation proof and a capability-named
trigger, but not the feature: tag and owning lane every other tracked growth in
this PR carries. It now names feature:host-substrate-measure-argument-surface and
v1-hand-queue-drain -- the same lane gunbc.target_invocation_seed_growth
owning_dissolution_lane names for this instrument's other hand-authored debt, so
the two halves of one seam are tracked in one place rather than two.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Sep 20, 2026
Merged via the queue into main with commit 75aaf98 Sep 20, 2026
4 checks passed
@gunbai-bot
gunbai-bot Bot deleted the session/proud-fox-12 branch September 20, 2026 15:00
gunbai-bot Bot pushed a commit that referenced this pull request Sep 20, 2026
…floor (§4c), so #11836's honest floor parses

floor_memory_demand (#11743) and namespace_reference_derived_residency_qualification (#11740) each carried a // block inside a fn body; the first honest queue run of this PR refused both at parse. Moved to module-item grain above their declaring fns.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
gunbai-bot Bot pushed a commit that referenced this pull request Sep 20, 2026
…ing, #11743)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
gunbai-bot Bot pushed a commit that referenced this pull request Sep 20, 2026
… to module grain (main's #11743 left it inside the declaration body, which refuses to parse and reds every floor run)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
gunbai-bot Bot pushed a commit that referenced this pull request Sep 20, 2026
…the parse phase stops failing and the floor can plan

DESIGN §4c admits only standalone leading `//` blocks attached to MODULE-SCOPE
declarations. Four files carry annotation lines indented inside a declaration
body, so `required-ci: phase parse` fails. The floor phase then refuses at
planning with ArmSetConsumerPlanningUnavailable — "no parse-phase declaration
index was lent to the floor" — selects no witness, and the job still concludes
success.

So every floor lane in the repository has been reporting a pass over an execution
that did not happen since #11743 and #11740 landed. The refusal is the consequence;
the parse failure is the earliest unjustified boundary (DESIGN §6b).

The population is the CENSUS, not the parse error output: the parse phase reported
thirteen lines in three files, but `grep -nE "^[[:space:]]+//"` over the corpus
finds a fourteenth. After this change that census returns zero files, which is the
stronger check because it does not depend on where the parser stopped.

Each block moves above the declaration it describes and is folded into that
declaration's existing leading annotation, with its subject named so the hoisted
text stays true of the whole declaration rather than of one arm:

  - floor_memory_demand.dag (191-194) — the unreadable-limit rationale joins the
    block above `qualify_floor_memory_demand`, stated as why that test comes
    before the three severity tests.
  - target_binding.dag (30-34) — the `FloorMemoryQualificationProducer` rationale
    joins the block above `type TargetProducer`, naming the variant it is about.
  - namespace_reference_derived_residency_qualification.dag (178-181) — a genuine
    BODY-POSITION comment inside a match arm, which §4c does not admit in any
    column, so unindenting in place would not have been the repair. It joins the
    block above `qualify_bounded_realization`, restated as the ordering of the
    whole fold rather than of the arm it sat in.
  - r2_permission_group_observe.dag (223) — one stray two-space `//` separator in
    an otherwise column-0 block.

No wording is dropped and no semantics change: annotations are erased before the
semantic passes, which is exactly why this is not a cosmetic edit — these were
failing at PARSE, so moving them changes parse success.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
gunbai-bot Bot pushed a commit that referenced this pull request Sep 20, 2026
…floor (§4c), so #11836's honest floor parses

floor_memory_demand (#11743) and namespace_reference_derived_residency_qualification (#11740) each carried a // block inside a fn body; the first honest queue run of this PR refused both at parse. Moved to module-item grain above their declaring fns.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit ae3519a)
gunbai-bot Bot pushed a commit that referenced this pull request Sep 20, 2026
…ing, #11743)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit c3891a6)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant