Skip to content

required floor is taking too long - #11837

Closed
briansrls wants to merge 9 commits into
mainfrom
fix-11743-fail-open
Closed

briansrls wants to merge 9 commits into
mainfrom
fix-11743-fail-open

Conversation

@briansrls

Copy link
Copy Markdown
Contributor

Auto-opened by session-dashboard for session nimble-wren-52.
Pushing to fix-11743-fail-open advances this PR.

Worker attestation

Before flipping this PR to ready for review, confirm each item:

  • Title describes the change (not the session id or branch).
  • PR body summarises what and why (replace the TODO below).
  • Tests run: name the command (e.g. npm test, cargo test) and the result.
  • If this closes a work item, the body contains a Closes #N directive.
  • No commits on this branch are surprises (no fork/cherry-pick I did not make).
  • No secrets / credentials / large binaries staged.

Summary

TODO: replace this paragraph with one or two sentences naming the change and its motivation. Reviewers read this first.

Test plan

  • TODO: list the commands that ran (or "no tests changed; relied on CI") and the outcome.

Brian Searls and others added 9 commits September 19, 2026 20:57
The passage quoted "22.74 GiB uncensored" as what the 2026-09-12 demand ruling
sized the runner slot from. The ruling says the opposite in as many words: it
records TWO disagreeing readings, takes the LARGER, and calls sizing to the
smaller "the fail-open direction". So this row attributed to the ruling the exact
arm the ruling refused, and named the fail-open direction as what the fail-closed
decision was.

The figure is REMOVED rather than corrected. This passage's own argument is that
the floor's demand does not apply to this deployment at all, and that holds at any
value -- so nothing downstream of it changes. It is a citation defect, not a wrong
decision. Carrying a second copy of a number owned by
gunbc.runner_slot_allocation gunbc_runner_slot_memory_max_ruling_note is the
DESIGN section 6 defect (name the producer, never transcribe its output), which is
also why no newer figure is substituted here even though one now exists.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ding

THE DEFECT THIS CLOSES. A peak read from a run that was HELD measures what the
process was ALLOWED, not what it needs. gunbc.runner_slot_allocation already
states that reading -- a peak equal to the throttle line is "the censored lower
bound ... the workload being held, not the workload being measured" -- and then,
on 2026-09-12, recorded two disagreeing figures for one run and said it did not
know which cgroup level each named. Nothing in the corpus made that ambiguity
unwritable, so the smaller figure propagated to gunbc.live_deploy.emit as "the
required floor's measured peak, uncensored".

gunbc.floor_memory_demand derives the standing from the kernel's own counters
instead of leaving it to prose. Three arms, and which three is the whole point:
DemandObserved (nothing held it -- the only arm citable as a demand),
DemandBounded (a lower bound, whether because the process was killed or because
it was throttled -- one arm because the consequence for a consumer is identical,
with the cause in the fields), and DemandUnreadable, which must never render like
either. A supervisor that cannot tell "could not read" from "it fit" reproduces
one layer out the conflation this exists to remove.

TWO AXES, KEPT APART. Termination (how the process ended, observed from outside)
and censoring (whether anything held it) are independent: a run can exit cleanly
having been throttled throughout, and a run can be killed having never been
throttled. Folding them into one enum loses exactly the pair an operator needs.

WHY THE SUBJECT IS THE PROCESS AND THE READING IS TAKEN FROM OUTSIDE. An
in-process observer -- a Drop guard, an exit hook, a final log line -- covers
ordinary returns and unwinding panics and nothing else. It does not run on
panic=abort, process::abort, process::exit or SIGKILL, and Rust's
allocation-error handler normally ABORTS. So allocation failure and the OOM kill,
the two terminations a memory instrument most exists to report, are precisely the
ones no in-process observer can report. The kernel maintains memory.peak and
memory.events regardless; a supervisor that outlived the child can still read
them. Measured 2026-09-19: a systemd-managed unit REAPS its cgroup on exit (the
peak is gone, though Result=oom-kill survives), while a cgroup the supervisor
created itself keeps both readable past a SIGKILL.

THE DISCRIMINATING PAIR IS TWO REAL RUNS, NOT TWO FIXTURES. Both on srv1 at
56375ec, same binary, same tree, back to back, one variable -- the throttle
line. Run A uncensored: peak 28962353152, all events zero, qualifies as a demand.
Run B at a 16 GiB line: peak 17181028352 with 17466 high events, does not. Four
further witnesses cover the cases counters alone would miss: a peak pinned
exactly to the line with zero events (the shape every figure behind the old 16/15
row had), a kill far below every limit, an unobserved termination, and an
unreadable cgroup. All six green by execution via claim_batch.

ALSO: one recurring_failure_mode row, filed from a near-miss I caused taking
these measurements. A cgroup placement write failed EPERM under a 2>/dev/null and
the workload ran UNCONSTRAINED on a shared host -- and an uncapped run is
indistinguishable from a capped one in everything the workload itself emits. The
row's general form is that a precondition establishing the ENVIRONMENT must be
verified by readback, never inferred from the call's exit status; the repo's own
ctrl-build "forwarding env: (none)" warning is the same class on a different knob.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ost arm pending)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…group read

THE INSTRUMENT. `gunbc test //gunbc/instruments:floor-memory-qualification` runs the
required-floor lane as a CHILD IN THE SUPERVISOR'S OWN CGROUP, waits, then reads
memory.peak and memory.events from outside the dead process and hands the raw
counters to `gunbc.floor_memory_demand` to judge. The three terminations map onto
the three arms: DemandObserved is 0, DemandBounded is 1 (a lower bound, which may
not size anything downward), every refusal is 2.

WHY THE CHILD SHARES THE CGROUP. Measured, not chosen: a systemd-managed unit REAPS
its cgroup on exit so the peak is gone before it can be read, and a self-created
cgroup cannot be JOINED by an existing process across a delegation boundary (EPERM).
Running the child in the cgroup the supervisor is already in avoids both -- nothing
is moved, and nothing reaps the cgroup because the supervisor still lives in it.

THE JUDGMENT STAYS IN THE SUBSTRATE. The host reads bytes and knows nothing about
what they mean; `qualify_floor_memory_from_readings` takes primitives and does every
interpretation in .dag over `extdeps.linux.cgroup_v2_memory`'s own parsers. A host
that built CgroupMemoryLimitValue itself would be a second parser for a file that
authority already owns.

TWO REFUSALS ADDED AFTER INVOKING IT EXPOSED THEM, both checked BEFORE the workload
so a 35-minute run is never spent producing an unattributable figure:

  MeasurementCgroupShared -- memory.peak is a property of the CGROUP, not a process.
  An ordinary login session scope was measured holding FIVE processes, so a bare
  invocation would report a neighbour's allocation as the floor's demand. Verified by
  execution: invoked in a session scope it refuses with exit 2 and names the pids.

  PeakDominatedByPriorHistory -- memory.peak is the cgroup's LIFETIME maximum and this
  kernel REFUSES to reset it (EPERM, measured). A post-run peak that did not rise above
  the pre-run baseline belongs to something that ran earlier, so it is refused rather
  than attributed to this run. The reset is attempted and READ BACK rather than trusted.

WHAT THE COMPILER DOES NOT CATCH, recorded because I claimed otherwise and was wrong:
there is no join between the .dag TargetProducer and the host's narrower Rust enum of
the same name, so adding a .dag variant forces NO host arm and the build stays green.
`RequiredFloorProducer` is the standing proof -- one occurrence corpus-wide, its own
declaration, no binding, no arm, no consumer. Registration was therefore verified by
INVOKING the label, not by compiling: the target now appears in `gunbc test`'s
available-target list and its refusal arm runs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ledger-Repair-Judged: docs/design-rung-drops.md
Ledger-Rows-Repaired: docs/design-rung-drops.md namespace_wave_admission_wall_removed
Heal-Candidate-Run: 35487819180
… rows nothing read

FINDING 1 (DESIGN section 7) — ACCEPTED AND FIXED. A 280-line hand-Rust host
module plus its host declarations landed with no row in
gunbc.target_invocation_seed_growth. The row now carries all nine
floor_memory_supervisor items, the two subject accessors and
run_floor_memory_qualification, and current_boundary names the new file and the
mod line.

ITS TRIGGER IS STATED SEPARATELY AND IS HONESTLY WEAKER than the rest of the row,
because assimilating it to the existing sentence would promise a discharge that is
not in sight. The other subsets dissolve when the emitter reaches their modeled
modules. This one does not: every operation in it is a resource effect — read
/proc/self/cgroup, spawn a child, wait on it, read /sys/fs/cgroup after the child
is gone — and resource operations resolve only inside a workflow function realized
by the seed INTERPRETER, which an emitted binary is not running under. It
discharges when the substrate can express a supervised child process as modeled
effects with typed refusals, and not before. Moving the judgment further into .dag
does not discharge it, and neither do witnesses over the read.

FINDING 2 (DESIGN section 3c) — ACCEPTED, AND THE FIX IS THE REVIEW'S SECOND
OPTION FOR A REASON I HAD TO MEASURE. floor_memory_qualification_source_roots and
floor_memory_qualification_lane had their only occurrences at their own
definitions, in the same file whose new comment states that exact test. The review
offered two remedies: route them through a real consumer, or drop them and let the
Rust own the fact explicitly.

I BUILT THE FIRST ONE, AND IT CORRUPTS THE MEASUREMENT. Reading those rows means
resolving a corpus graph, and this supervisor shares its cgroup with the child BY
DESIGN — that sharing is what lets memory.peak survive the child's death. So the
resolve lands in the very counter the instrument reports. Same failing floor, same
tree, same binary, differing only in whether the subject was read from the model:

  Rust-owned subject   peak 15746146304  (14.66 GiB)
  Rust-owned subject   peak 15704227840  (14.63 GiB)
  read from the model  peak 22293544960  (20.76 GiB)   <- +6.1 GiB, 42% inflation
  Rust-owned again     peak 15816912896  (14.73 GiB)   <- restored

An instrument may not consult the authority from inside the cgroup it measures: the
act of reading perturbs the reading. Netting the supervisor's footprint back out
was not available either — that replaces a measured number with an adjusted one,
which is the habit gunbc.floor_memory_demand exists to refuse.

So the two rows are DELETED rather than left unconsumed, and the Rust says plainly
that it owns the subject and why, with the figures above and with the condition
that would restore the modeled form: any route that reads them OUTSIDE the measured
cgroup. That needs the cgroup lifecycle modeled, which is the same capability the
seed-growth trigger names.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…e fail-open

FOUR FINDINGS, ALL VERIFIED AGAINST THE CODE BEFORE FIXING.

1. THE CONTAMINATION GUARD FAILED OPEN ON THE CASE IT WAS BUILT FOR, and this is
the serious one. `cgroup.procs` lists DIRECT members; `memory.peak` aggregates the
whole SUBTREE. `nearest_cgroup_with_peak` deliberately climbs to an ancestor, so
the guard could find no strangers in a cgroup whose descendants were running
anything at all. Measured on srv1 while confirming it: user-1000.slice has ZERO
direct processes — a direct-membership check finds nothing — against 242 child
cgroups, 336 processes beneath it, and memory.peak 392042180608. The instrument
would have reported a third of a terabyte of co-tenant allocation as the floor's
demand, as DemandObserved.

Fixed twice over: membership is now walked over the SUBTREE, and a cgroup with any
child at all is refused separately, because a descendant can be created after the
check and only a leaf is stable. EXERCISED, not asserted: in a delegated scope with
a deliberately created child cgroup the instrument refuses with
MeasurementCgroupHasChildren and exit 2.

2. THE REFUSAL VOCABULARY HAD FORKED IN BOTH DIRECTIONS. `DemandReadRefusalCause`
carried three arms nothing could construct (a check whose forbidden state is
unwritable is a decoration, section 4b) while the producer carried five real causes
the model had never heard of — and every refusal the instrument actually emits came
from the unmodeled set.

The repair is not to copy one list into the other: THE TWO POPULATIONS HAVE
DIFFERENT SUBJECTS. A supervisor's refusals happen BEFORE there is anything to
judge — no cgroup, no child, no attributable counter — and terminate the invocation
with no observation, never reaching the fold. What reaches the judgment is a
complete set of readings, so the only way IT can refuse is that a reading cannot be
interpreted. One arm, because there is exactly one such way. The Rust comment
claiming to mirror the model is corrected to say the opposite and why.

`TerminationUnobserved` is kept and its reachability stated plainly: exercised by
the witness, not by today's producer, because the fail-closed reading it encodes is
a property of the judgment rather than of one producer.

3. THE SEED CENSUS WAS ITEM-INCOMPLETE — five private helpers omitted, including
`nearest_cgroup_with_peak`, which carries the ancestor-climb decision behind finding
1. Privacy is not the grain this roster uses; it already enumerates private helpers
of the sibling module. Six added.

4. `peak_bytes: Int` — two reviewers split on this row (68906 refuted it as a
correct host-to-substrate primitive boundary, 68936 echoed it as unrefuted). It now
carries an explicit dissolve-on gate rather than an argument, stating why the
parameter is primitive at this one seam, that the scalar does not propagate, and
what would dissolve it: an interpreter argument surface that admits a constructed
ByteSize. Explicitly NOT dissolved by wrapping the literal one frame outward.

A new witness drives the refusal through the real entry point rather than
hand-constructing the variant, so the arm is evidence about the production route.
All seven witnesses green by execution.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 20, 2026

Copy link
Copy Markdown
Contributor

Closing: this is an auto-opened placeholder on a throwaway branch I created while investigating review 68949 on #11743. It carries no work of its own — its commits are proud-fox-12's, and that lane fixed the finding itself on session/proud-fox-12. Nothing here is unlanded and nothing should be merged from it.

— sent from nimble-wren-52

@gunbai-bot gunbai-bot Bot closed this Sep 20, 2026
@gunbai-bot
gunbai-bot Bot deleted the fix-11743-fail-open branch September 20, 2026 12:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant