Repository navigation
Conversation
The passage quoted "22.74 GiB uncensored" as what the 2026-09-12 demand ruling sized the runner slot from. The ruling says the opposite in as many words: it records TWO disagreeing readings, takes the LARGER, and calls sizing to the smaller "the fail-open direction". So this row attributed to the ruling the exact arm the ruling refused, and named the fail-open direction as what the fail-closed decision was. The figure is REMOVED rather than corrected. This passage's own argument is that the floor's demand does not apply to this deployment at all, and that holds at any value -- so nothing downstream of it changes. It is a citation defect, not a wrong decision. Carrying a second copy of a number owned by gunbc.runner_slot_allocation gunbc_runner_slot_memory_max_ruling_note is the DESIGN section 6 defect (name the producer, never transcribe its output), which is also why no newer figure is substituted here even though one now exists. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ding THE DEFECT THIS CLOSES. A peak read from a run that was HELD measures what the process was ALLOWED, not what it needs. gunbc.runner_slot_allocation already states that reading -- a peak equal to the throttle line is "the censored lower bound ... the workload being held, not the workload being measured" -- and then, on 2026-09-12, recorded two disagreeing figures for one run and said it did not know which cgroup level each named. Nothing in the corpus made that ambiguity unwritable, so the smaller figure propagated to gunbc.live_deploy.emit as "the required floor's measured peak, uncensored". gunbc.floor_memory_demand derives the standing from the kernel's own counters instead of leaving it to prose. Three arms, and which three is the whole point: DemandObserved (nothing held it -- the only arm citable as a demand), DemandBounded (a lower bound, whether because the process was killed or because it was throttled -- one arm because the consequence for a consumer is identical, with the cause in the fields), and DemandUnreadable, which must never render like either. A supervisor that cannot tell "could not read" from "it fit" reproduces one layer out the conflation this exists to remove. TWO AXES, KEPT APART. Termination (how the process ended, observed from outside) and censoring (whether anything held it) are independent: a run can exit cleanly having been throttled throughout, and a run can be killed having never been throttled. Folding them into one enum loses exactly the pair an operator needs. WHY THE SUBJECT IS THE PROCESS AND THE READING IS TAKEN FROM OUTSIDE. An in-process observer -- a Drop guard, an exit hook, a final log line -- covers ordinary returns and unwinding panics and nothing else. It does not run on panic=abort, process::abort, process::exit or SIGKILL, and Rust's allocation-error handler normally ABORTS. So allocation failure and the OOM kill, the two terminations a memory instrument most exists to report, are precisely the ones no in-process observer can report. The kernel maintains memory.peak and memory.events regardless; a supervisor that outlived the child can still read them. Measured 2026-09-19: a systemd-managed unit REAPS its cgroup on exit (the peak is gone, though Result=oom-kill survives), while a cgroup the supervisor created itself keeps both readable past a SIGKILL. THE DISCRIMINATING PAIR IS TWO REAL RUNS, NOT TWO FIXTURES. Both on srv1 at 56375ec, same binary, same tree, back to back, one variable -- the throttle line. Run A uncensored: peak 28962353152, all events zero, qualifies as a demand. Run B at a 16 GiB line: peak 17181028352 with 17466 high events, does not. Four further witnesses cover the cases counters alone would miss: a peak pinned exactly to the line with zero events (the shape every figure behind the old 16/15 row had), a kill far below every limit, an unobserved termination, and an unreadable cgroup. All six green by execution via claim_batch. ALSO: one recurring_failure_mode row, filed from a near-miss I caused taking these measurements. A cgroup placement write failed EPERM under a 2>/dev/null and the workload ran UNCONSTRAINED on a shared host -- and an uncapped run is indistinguishable from a capped one in everything the workload itself emits. The row's general form is that a precondition establishing the ENVIRONMENT must be verified by readback, never inferred from the call's exit status; the repo's own ctrl-build "forwarding env: (none)" warning is the same class on a different knob. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ost arm pending) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…group read THE INSTRUMENT. `gunbc test //gunbc/instruments:floor-memory-qualification` runs the required-floor lane as a CHILD IN THE SUPERVISOR'S OWN CGROUP, waits, then reads memory.peak and memory.events from outside the dead process and hands the raw counters to `gunbc.floor_memory_demand` to judge. The three terminations map onto the three arms: DemandObserved is 0, DemandBounded is 1 (a lower bound, which may not size anything downward), every refusal is 2. WHY THE CHILD SHARES THE CGROUP. Measured, not chosen: a systemd-managed unit REAPS its cgroup on exit so the peak is gone before it can be read, and a self-created cgroup cannot be JOINED by an existing process across a delegation boundary (EPERM). Running the child in the cgroup the supervisor is already in avoids both -- nothing is moved, and nothing reaps the cgroup because the supervisor still lives in it. THE JUDGMENT STAYS IN THE SUBSTRATE. The host reads bytes and knows nothing about what they mean; `qualify_floor_memory_from_readings` takes primitives and does every interpretation in .dag over `extdeps.linux.cgroup_v2_memory`'s own parsers. A host that built CgroupMemoryLimitValue itself would be a second parser for a file that authority already owns. TWO REFUSALS ADDED AFTER INVOKING IT EXPOSED THEM, both checked BEFORE the workload so a 35-minute run is never spent producing an unattributable figure: MeasurementCgroupShared -- memory.peak is a property of the CGROUP, not a process. An ordinary login session scope was measured holding FIVE processes, so a bare invocation would report a neighbour's allocation as the floor's demand. Verified by execution: invoked in a session scope it refuses with exit 2 and names the pids. PeakDominatedByPriorHistory -- memory.peak is the cgroup's LIFETIME maximum and this kernel REFUSES to reset it (EPERM, measured). A post-run peak that did not rise above the pre-run baseline belongs to something that ran earlier, so it is refused rather than attributed to this run. The reset is attempted and READ BACK rather than trusted. WHAT THE COMPILER DOES NOT CATCH, recorded because I claimed otherwise and was wrong: there is no join between the .dag TargetProducer and the host's narrower Rust enum of the same name, so adding a .dag variant forces NO host arm and the build stays green. `RequiredFloorProducer` is the standing proof -- one occurrence corpus-wide, its own declaration, no binding, no arm, no consumer. Registration was therefore verified by INVOKING the label, not by compiling: the target now appears in `gunbc test`'s available-target list and its refusal arm runs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ledger-Repair-Judged: docs/design-rung-drops.md Ledger-Rows-Repaired: docs/design-rung-drops.md namespace_wave_admission_wall_removed Heal-Candidate-Run: 35487819180
… rows nothing read FINDING 1 (DESIGN section 7) — ACCEPTED AND FIXED. A 280-line hand-Rust host module plus its host declarations landed with no row in gunbc.target_invocation_seed_growth. The row now carries all nine floor_memory_supervisor items, the two subject accessors and run_floor_memory_qualification, and current_boundary names the new file and the mod line. ITS TRIGGER IS STATED SEPARATELY AND IS HONESTLY WEAKER than the rest of the row, because assimilating it to the existing sentence would promise a discharge that is not in sight. The other subsets dissolve when the emitter reaches their modeled modules. This one does not: every operation in it is a resource effect — read /proc/self/cgroup, spawn a child, wait on it, read /sys/fs/cgroup after the child is gone — and resource operations resolve only inside a workflow function realized by the seed INTERPRETER, which an emitted binary is not running under. It discharges when the substrate can express a supervised child process as modeled effects with typed refusals, and not before. Moving the judgment further into .dag does not discharge it, and neither do witnesses over the read. FINDING 2 (DESIGN section 3c) — ACCEPTED, AND THE FIX IS THE REVIEW'S SECOND OPTION FOR A REASON I HAD TO MEASURE. floor_memory_qualification_source_roots and floor_memory_qualification_lane had their only occurrences at their own definitions, in the same file whose new comment states that exact test. The review offered two remedies: route them through a real consumer, or drop them and let the Rust own the fact explicitly. I BUILT THE FIRST ONE, AND IT CORRUPTS THE MEASUREMENT. Reading those rows means resolving a corpus graph, and this supervisor shares its cgroup with the child BY DESIGN — that sharing is what lets memory.peak survive the child's death. So the resolve lands in the very counter the instrument reports. Same failing floor, same tree, same binary, differing only in whether the subject was read from the model: Rust-owned subject peak 15746146304 (14.66 GiB) Rust-owned subject peak 15704227840 (14.63 GiB) read from the model peak 22293544960 (20.76 GiB) <- +6.1 GiB, 42% inflation Rust-owned again peak 15816912896 (14.73 GiB) <- restored An instrument may not consult the authority from inside the cgroup it measures: the act of reading perturbs the reading. Netting the supervisor's footprint back out was not available either — that replaces a measured number with an adjusted one, which is the habit gunbc.floor_memory_demand exists to refuse. So the two rows are DELETED rather than left unconsumed, and the Rust says plainly that it owns the subject and why, with the figures above and with the condition that would restore the modeled form: any route that reads them OUTSIDE the measured cgroup. That needs the cgroup lifecycle modeled, which is the same capability the seed-growth trigger names. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
# Conflicts: # docs/design-rung-drops.md
…e fail-open FOUR FINDINGS, ALL VERIFIED AGAINST THE CODE BEFORE FIXING. 1. THE CONTAMINATION GUARD FAILED OPEN ON THE CASE IT WAS BUILT FOR, and this is the serious one. `cgroup.procs` lists DIRECT members; `memory.peak` aggregates the whole SUBTREE. `nearest_cgroup_with_peak` deliberately climbs to an ancestor, so the guard could find no strangers in a cgroup whose descendants were running anything at all. Measured on srv1 while confirming it: user-1000.slice has ZERO direct processes — a direct-membership check finds nothing — against 242 child cgroups, 336 processes beneath it, and memory.peak 392042180608. The instrument would have reported a third of a terabyte of co-tenant allocation as the floor's demand, as DemandObserved. Fixed twice over: membership is now walked over the SUBTREE, and a cgroup with any child at all is refused separately, because a descendant can be created after the check and only a leaf is stable. EXERCISED, not asserted: in a delegated scope with a deliberately created child cgroup the instrument refuses with MeasurementCgroupHasChildren and exit 2. 2. THE REFUSAL VOCABULARY HAD FORKED IN BOTH DIRECTIONS. `DemandReadRefusalCause` carried three arms nothing could construct (a check whose forbidden state is unwritable is a decoration, section 4b) while the producer carried five real causes the model had never heard of — and every refusal the instrument actually emits came from the unmodeled set. The repair is not to copy one list into the other: THE TWO POPULATIONS HAVE DIFFERENT SUBJECTS. A supervisor's refusals happen BEFORE there is anything to judge — no cgroup, no child, no attributable counter — and terminate the invocation with no observation, never reaching the fold. What reaches the judgment is a complete set of readings, so the only way IT can refuse is that a reading cannot be interpreted. One arm, because there is exactly one such way. The Rust comment claiming to mirror the model is corrected to say the opposite and why. `TerminationUnobserved` is kept and its reachability stated plainly: exercised by the witness, not by today's producer, because the fail-closed reading it encodes is a property of the judgment rather than of one producer. 3. THE SEED CENSUS WAS ITEM-INCOMPLETE — five private helpers omitted, including `nearest_cgroup_with_peak`, which carries the ancestor-climb decision behind finding 1. Privacy is not the grain this roster uses; it already enumerates private helpers of the sibling module. Six added. 4. `peak_bytes: Int` — two reviewers split on this row (68906 refuted it as a correct host-to-substrate primitive boundary, 68936 echoed it as unrefuted). It now carries an explicit dissolve-on gate rather than an argument, stating why the parameter is primitive at this one seam, that the scalar does not propagate, and what would dissolve it: an interpreter argument surface that admits a constructed ByteSize. Explicitly NOT dissolved by wrapping the literal one frame outward. A new witness drives the refusal through the real entry point rather than hand-constructing the variant, so the arm is evidence about the production route. All seven witnesses green by execution. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Contributor
|
Closing: this is an auto-opened placeholder on a throwaway branch I created while investigating review 68949 on #11743. It carries no work of its own — its commits are proud-fox-12's, and that lane fixed the finding itself on — sent from nimble-wren-52 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Auto-opened by session-dashboard for session
nimble-wren-52.Pushing to
fix-11743-fail-openadvances this PR.Worker attestation
Before flipping this PR to ready for review, confirm each item:
npm test,cargo test) and the result.Closes #Ndirective.Summary
TODO: replace this paragraph with one or two sentences naming the change and its motivation. Reviewers read this first.
Test plan