Skip to content

Ground systemd slice directives on the property authority: five knobs, zero new names - #9552

Merged
briansrls merged 3 commits into
mainfrom
session/smart-pike-831
Aug 28, 2026
Merged

briansrls merged 3 commits into
mainfrom
session/smart-pike-831

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Aug 28, 2026

Copy link
Copy Markdown
Contributor

What this changes

extdeps.systemd.unit_file SystemdSliceDirective was a two-arm coproduct — SliceMemoryMax { bytes: NonEmptyStr } and SliceMemoryHigh { bytes: NonEmptyStr }. The fabric cell resource boundary (gunbc.fabric_cell_converge fabric_cell_resource_boundary_state_digest) declares five values, so realizing it needs MemorySwapMax=, TasksMax= and CPUWeight= as well.

Adding three arms was the obvious move and is the wrong one: it spells knob names that extdeps.systemd SystemdUnitProperty already carries, widening the fork gunbc.systemd_property_directive_overlap counts from two names to five.

So the directive is now a sole_constructor record pairing the property with a modeled value, reachable only through per-knob mints:

type SystemdSliceDirective sole_constructor {
  property: SystemdUnitProperty
  value: SystemdDirectiveValue
}

Zero knob names are minted. The wire spelling has exactly one owner — systemd_unit_property_wire, reached through property — so systemd_slice_directive_line has no match left to get wrong, and a sixth knob is a mint with no edit to the renderer.

Why the show surface is safe here

gunbc.systemd_property_directive_overlap rules that SystemdUnitProperty is the show surface: it mixes settable knobs with observations the manager computes, so a directive type admitting any member would make MemoryCurrent=, ActiveState= and MainPID= writable into a unit file.

What closes that is not a check and not a narrower field type. The only constructors in existence are the mints, and sole_constructor refuses a record literal from any other module. MemoryCurrent= has no spelling because nothing produces it.

Stated honestly: "no observation is writable" is enforced by the mint set plus sole_constructor, not by a type-level refusal. Same-module construction remains possible by design — the confinement is about who else may write the pairing.

Scope of the sole_constructor guarantee (declared, not inherited)

sole_constructor confines cross-module construction on the source→.dag path only. DESIGN records by execution that an emitted Rust mirror of such a type is silently forgeable (extdeps.uri UriValidatedScalar). This record is constructed and rendered entirely inside the .dag pipeline — never deserialized, nothing reconstructs it on the emitted side — so the mint is sufficient for the path it travels. The carrier says so, so the next person gets a fresh question rather than an inherited guarantee.

Two deviations from the brief (both approved by silent-bear-842)

  1. Nine mints, not five. CPUWeight gets no unlimited mint — systemd defines no such value for it. The other four knobs each get one. The asymmetry is the model.
  2. Four value arms, not three. DirectiveCardinal (TasksMax, a count of processes) is split from DirectiveWeight (CPUWeight, a relative share against sibling cgroups). They render identically today, but a count and a share are different quantities; fusing them puts the difference into position. Identical rendering is not proven coincidence of the quantities.

Evidence — mutation receipt, not just a green roster

Compile over the witness closure: 0 blocking, 205 advisory (corpus-normal).

witness baseline mint mis-wired restored
each_slice_knob_renders_its_own_systemd_name true false true
the_five_slice_knobs_render_five_distinct_names true false —
a_slice_unit_file_serializes_to_exactly_these_bytes true false true

The mutation wired slice_memory_high to MemoryMax — the defect this shape actually admits, which compiles clean and renders a plausible line while writing the wrong ceiling onto a host. Tree restored byte-exactly (git diff empty) and the green re-confirmed; a restore that isn't re-verified is an assumption.

The byte-exact slice unit test passing unchanged is the load-bearing one: the carrier swap preserved the deployed bytes.

Also enrolled: the four sentinel renderings, and an_unlimited_ceiling_is_not_a_zero_ceiling — because a DirectiveUnlimited arm collapsing into the numeric one writes a zero ceiling, which is not "no limit" but the strictest limit there is.

Name-collision sweep

Two declarations of one spelling resolve per-entry and refuse whole-corpus, reporting at an innocent third module — and an entry-scoped witness run structurally cannot see it. All 20 names this branch adds were swept against dag/ + src/v2 on two independent passes (declaration-form regex, whole-word grep): zero collisions. Sweep proven non-blind first — NotReady and MemoryMax return hits under the same regex. The value arms are Directive-prefixed rather than bare Unlimited/Count/Weight, which is where the tempting short names sit.

The multi-part row was rewritten, not deleted

gunbc.systemd_directive_value_grain's dissolution was premised on SliceMemoryMax surviving as a named declaration; this cut falsifies that. The row carries three obligations — slice value grain, Environment becoming two fields, the sccache name move — and this cut discharges only the first. Only the falsified clause moved; Environment and the sccache move stand verbatim and still gate deletion. systemd_property_directive_overlap is named as the reason the shape moved.

Residue kept honest: the percentage form (MemoryMax=50%) remains unmodeled and is still filed. The byte count and the suffixed form are one ByteSize and no longer a lost distinction; the percentage is a different quantity and is not claimed as discharged. TasksMax's bare Int is a second residue, already declared upstream against feature:task-count-measure-carrier.

What this does not do

It unblocks the cell resource boundary; it does not realize it. Nothing in the fabric-cell family renders a .slice unit file today — the only SliceUnitFile producer is gunbc.build_cache_unit. The realization lives on fabric/cell-realization (#9467), unmerged, and is not this cut's.

🤖 Generated with Claude Code

https://claude.ai/code/session_01G2LsQXUHjevegFefaPU9BE

Brian Searls and others added 3 commits August 28, 2026 01:03
…, zero new names

SystemdSliceDirective was a two-arm coproduct spelling MemoryMax and MemoryHigh a
second time and carrying both values as NonEmptyStr. The fabric cell resource
boundary declares five values, so realizing it needed MemorySwapMax=, TasksMax=
and CPUWeight= -- and adding three arms would have widened the knob-name fork
gunbc.systemd_property_directive_overlap counts from two to five.

Instead the directive is a sole_constructor record pairing a SystemdUnitProperty
with a modeled SystemdDirectiveValue, reachable only through per-knob mints. Zero
knob names are minted: the wire spelling has one owner, systemd_unit_property_wire,
reached through the property. MemoryCurrent= has no spelling because no mint
produces it, not because a check rejects it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G2LsQXUHjevegFefaPU9BE
@gunbai-bot

gunbai-bot Bot commented Aug 28, 2026

Copy link
Copy Markdown
Contributor Author

Why this PR is red, and why no fix is owed from it

The two failing checks are one cause counted twice. At 6695d62:

check result
required-witnesses-build success
required-witnesses-floor failure
witnesses failure — the aggregator, red whenever either lane is red

witnesses is not an independent defect; it needs both lanes and reports their conjunction.

The floor failure is inherited, and the counts say so

required-floor: planned=11999 executed=11999 terminal=11999 passed=11856 failed=0
                completed_over_cost_requirement=1
required-floor: verdict=FloorRefused unexpected_failures=0
COMPLETED-OVER-COST-REQUIREMENT
  test.claim.callable_candidate_ambiguity_witness.neither_green_source_refuses_and_neither_mis_resolves
  exceeded its budget: Cpu, cost exactly 5058ms against 5000ms

failed=0 and unexpected_failures=0: nothing failed. One claim PASSED and then exceeded a Cpu budget, and the floor refuses on cost rather than on a defect. It is the same claim, with the same cause, on main and on every PR based on it. No version of this diff clears it — the remedy is #9477 (first-payer billing), owned elsewhere and under hold.

The build red was inherited too, and is now gone

It was regen FAIL generated surface drift in four emitter mirrors — identical on main at a commit that is an ancestor of this branch. I declined to regenerate those files: two byte-identical fixes already existed (#9537, #9551), so a third would have been a triplicate of a change I did not author, and greening this PR over a break still live on main would have let the next reader take a green required check as evidence the tree was clean. Merging #9551 cleared it with no change to this branch's four files.

This PR's own subject is verified

  • compile over the witness closure: 0 blocking
  • floor: failed=0; roster planned 11996 → 11999, exactly this branch's net +3 witnesses
  • mutation receipt: wiring slice_memory_high to MemoryMax flips three witnesses to false; tree restored byte-exactly; green returns

One correction to my own earlier reporting

I described the cost row as "real rather than marginal" from two observations. A third makes that wrong: 5812ms (16.2% over) → 5535ms (10.7%) → 5058ms (1.2%). Reproducible holds; not marginal does not. It may go green on a quiet runner, which would make this an intermittently red gate — so a single passing run should not be read as fixed. Three uncontrolled points are not a distribution, and that caveat applies to this reading too.

— sent from smart-pike-831

@gunbai-bot

gunbai-bot Bot commented Aug 28, 2026

Copy link
Copy Markdown
Contributor Author

Correction to the cost-row reading above: it is not monotone

A fourth observation (#9467, floor run 33136317264) breaks the sequence I described. Ordered by time, against the 5000ms Cpu gate:

cost over subject
5812ms 16.2% main
5535ms 10.7% #9552
5553ms 11.1% #9467 — up from the previous
5058ms 1.2% #9552 @ 6695d62

My comment above read this as monotonically decreasing. It is not: 5535 → 5553 goes the wrong way, and the spread across four points is 754ms.

The conclusion is unchanged and better supported. A monotone sequence would have been consistent with a trend — something changing across the commit sequence. Non-monotone with a 754ms spread is what load noise looks like, which is the reading the correction moved to. So the wrong half was my stated evidence, not the verdict.

What this sharpens: if the spread is noise around a value near the line, 5058ms being 1.2% over means a quiet runner plausibly lands under 5000. That makes this an intermittently red required gate — worse than a stably red one, because it manufactures false attributions: whoever runs while it is over the line will go looking in their own diff.

What is still not established, so four points are not mistaken for a distribution: no control for runner identity, no control for concurrent load, no repeated measurement on a single commit. Four points from four commits on four runner acquisitions cannot separate load from change. Measuring one commit N times on one runner is what would settle it.

— sent from smart-pike-831

@gunbai-bot

gunbai-bot Bot commented Aug 28, 2026

Copy link
Copy Markdown
Contributor Author

Cost row, settled — this supersedes the reasoning in my two comments above

Both earlier comments reasoned from the gated row's cost in isolation (first "not marginal", then "monotone decreasing"). Both readings were unnecessary: the module has sibling rows in the same entry, and every floor log already prints their costs, so each run carries its own internal control.

run / slot gated sibling ratio
main / srv3-05 5812 3130 1.857
#9552 / srv4-05 5535 3068 1.804
6695d62 / srv2-04 5058 2833 1.785
#9467 / srv2-03 5049 2746 1.839

Gated varies 15.1%. Sibling varies 14.0%. Their ratio varies 4.0%. A third sibling corroborates: it drops off the over-cost list entirely on the fast runs, scaling the same ~0.92.

What the ratio rules out. First-toucher attribution moves a fixed shared charge between siblings, so one falls as another jumps and the ratio breaks. Real cost drift moves the gated row against its siblings. Neither happened. A whole-entry scaling factor — host speed and load — is what survives.

Independently supporting this, from the same run: cache=module_path_index fill_ms=60527 paid_by=<outside-fold> consumer_claims=335 — 60.5s of shared fill consumed by 335 claims including this one and explicitly not billed to it.

The consequence, which inverts the natural instinct

  • a 1.0% cut clears the 5000ms gate on the fastest host observed
  • a 14.0% cut clears it on the slowest

So any reduction between 1% and 14% converts a mostly-red gate into a permanently intermittent one — green on srv2, red on srv3. A small successful optimisation makes this harder to diagnose, not easier, because an intermittent gate manufactures false attributions: whoever runs while it is over the line goes hunting in their own diff. Deterministic green must beat the slowest host with headroom.

Limits, and how to refute this

n=4, every point a different commit on a different slot. I did not control either — that is what the normalizer is for — but it means no distribution and no flip probability. A pinned-runner campaign was authorised and then cancelled as not worth the CI while the remedy is under hold.

Refutation is one grep: a run where the gated row moves and its siblings do not.

None of this is actionable in this PR. The remedy is #9477, owned elsewhere and under hold.

— sent from smart-pike-831

@gunbai-bot

gunbai-bot Bot commented Aug 28, 2026

Copy link
Copy Markdown
Contributor Author

Correction — a fifth reading partly refutes the model in my previous comment, and moves the number

A fifth observation (main 5f782fb6e, slot srv1-09) at 6889ms, with its siblings:

run / slot gated sibling ratio
main / srv3-05 5812 3130 1.857
#9552 / srv4-05 5535 3068 1.804
6695d62 / srv2-04 5058 2833 1.785
#9467 / srv2-03 5049 2746 1.839
main / srv1-09 6889 3406 2.023

What breaks. I reported the ratio stable at 4.0% and argued from it that the variance was purely whole-entry host scaling. Across five points the spread is 13.3%, and the new point is the outlier. Tested as a prediction rather than by eye: the model (gated ≈ 1.82 × sibling) predicts 6199ms at sibling=3406; observed 6889 — a 10% miss, in the direction that matters. Gated varies 36.4% across five runs while its sibling varies 24.0%, so the gated row moves more than its sibling.

That is the weaker form of the refutation criterion I published above. I wrote "a run where the gated row moves and its siblings do not"; this is "moves disproportionately." I did not anticipate that form, and it still counts.

What survives. All three rows still move together directionally — srv1-09 is the maximum for every one of them. Host speed is real and still dominant. What is dead is the claim that it is the only term and that the row has a stable intrinsic cost. There is a residual hitting the gated row harder, and I do not know what it is.

The number changes, and this is the part that matters downstream:

  • clearing 5000 on the fastest observed (5049) still needs 1.0%
  • clearing 5000 on the slowest observed (6889) now needs 27.4% — I previously said 14.0%

So the trap is wider than I reported: any cut between 1% and 27.4% leaves a permanently intermittent gate. The shape is unchanged; the magnitude nearly doubled. And with the ratio no longer stable, 27.4% is a lower bound, not a ceiling — a slower host moves it again, and my model should not be used to extrapolate, because it just failed a prediction by 10%.

The open question, now specific: why does this row scale worse than its siblings on a slow host, when a pure load factor would move them together?

Still nothing actionable in this PR; the remedy is #9477, under hold.

— sent from smart-pike-831

@gunbai-bot

gunbai-bot Bot commented Aug 28, 2026

Copy link
Copy Markdown
Contributor Author

Two label corrections to the table above (neither changes the conclusion)

1. Five runs, four subjects — not five independent subjects. Two rows are the same PR at two heads: 0aae94b (5535) and 6695d62 (5058) are both this branch, the first an ancestor of the second. Corrected labels:

run subject slot gated sibling ratio
main b6003a45e main srv3-05 5812 3130 1.857
#9552 @ 0aae94b this PR srv4-05 5535 3068 1.804
#9552 @ 6695d62 this PR srv2-04 5058 2833 1.785
#9467 other PR srv2-03 5049 2746 1.839
main 5f782fb6e main srv1-09 6889 3406 2.023

The host-variance reading is unaffected — five runs on five different slots is still five samples of the confound — but the points are less independent as subjects than a quick read of the earlier table suggests, and that is worth knowing if the table travels.

2. The prediction error is 10.0% or 11.1% depending on the denominator, and I should have named which I used. 690ms gap: 10.0% relative to the observed value (the usual convention for prediction error, and what I reported), 11.1% relative to the predicted value. Either way the model missed and the refutation stands; only the reported figure needs a label.

A sixth reading exists and is not in this table: #9553 at 5397ms, siblings unfetched. Full gated set: 5049 / 5058 / 5397 / 5535 / 5812 / 6889. I have stopped adding uncontrolled points — they move the number without explaining it — so it is recorded here rather than analysed.

The live unknown, unassigned: why does this row scale worse than its siblings on a slow host, when a pure load factor would move them together?

— sent from smart-pike-831

@briansrls
briansrls merged commit 3a8344b into main Aug 28, 2026
1 of 3 checks passed
@briansrls
briansrls deleted the session/smart-pike-831 branch August 28, 2026 05:32
@briansrls
briansrls restored the session/smart-pike-831 branch August 28, 2026 05:39
@gunbai-bot

gunbai-bot Bot commented Aug 28, 2026

Copy link
Copy Markdown
Contributor Author

Final correction: a same-commit re-run refutes my model outright, and excludes the tree

I re-ran the floor on this exact head (no new commit, so the approvals stand). It refused again — 5883ms, same identity, failed=0. One deliberate attempt; not repeating it.

The re-run is the controlled measurement, because it replays the same commit. This is the only pair in seven readings where the tree is held fixed:

commit slot gated sibling ratio
6695d62 srv2-04 5058 2833 1.785
6695d62 srv3-15 5883 3064 1.920

Identical input. Gated moved +16.3%; its sibling moved +8.2%.

  1. The tree is excluded as a cause. Same bytes, 16.3% apart. No commit-sequence or cost-drift explanation survives this.
  2. My scaling model is refuted outright, not partly. A whole-entry load factor requires both rows to move by the same proportion; the gated row moved twice as far as its control. Predicting from the sibling at the old ratio gives 5469 vs 5883 observed. The earlier comment called this partial, on one outlier — on a controlled pair it is not partial. Do not extrapolate with that ratio, including the 27.4% figure, which assumed it.
  3. The open question changes shape. I asked why this row scales worse than its siblings on a slow host. Wrong framing — it does so on the same tree across two ordinary hosts, neither being the slow outlier. The row has environment-sensitivity its siblings do not share. That is a property of the row, not of the fleet.

What is reliable after seven readings (5049 / 5058 / 5397 / 5535 / 5812 / 5883 / 6889, six over the gate): the identity is stable across every run, the observed range is 5049–6889, the tree is excluded, and this row responds to environment roughly twice as strongly as a sibling in the same entry.

Nothing here is actionable in this PR. Remedy is #9477, under hold.

— sent from smart-pike-831

@gunbai-bot

gunbai-bot Bot commented Aug 28, 2026

Copy link
Copy Markdown
Contributor Author

Withdrawing the withdrawal: 27.4% never used the refuted model

My previous comment said "do not extrapolate with that ratio, including the 27.4% figure, which assumed it." That clause is wrong and I am retracting it.

27.4% is (6889 − 5000) / 6889 — arithmetic on one observed reading and the gate. No sibling, no ratio, no prediction enters it. My model dying does not touch it.

I conflated two different things: the revision from 14% to 27.4% came from a new observation (correct), and the model justified the revision (false — the model only ever justified the stability and extrapolation claims). Over-withdrawing a sound number puts a false record into circulation exactly as under-withdrawing a bad one does.

The accurate split after seven readings:

Reliable — all observations, no fit identity stable across all seven runs · observed range 5049–6889 · clearing the gate on the worst observed run requires 27.4% · the tree is excluded as a cause (same-commit pair) · the row is ~2× as environment-sensitive as its sibling
Dead the ratio model · any extrapolation beyond the observed range, including any claim that 27.4% is sufficient

"Lower bound, not a ceiling" remains the right framing, and it was always justified by the existence of a slower host rather than by the ratio.

Corrections on this thread now stand at: the marginal/real call, the monotone claim, the row labels, the error denominator, the model itself — and this one, which is a correction to a correction. Each changed the record; none changed the PR, which carries no defect any of these runs found.

— sent from smart-pike-831

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant