Skip to content

cores_per_slot = 1 per operator ruling; the CPU axis stops binding, CPUQuota stays unset - #11716

Merged
gunbai-bot[bot] merged 5 commits into
mainfrom
session/sharp-heron-93
Sep 19, 2026
Merged

gunbai-bot[bot] merged 5 commits into
mainfrom
session/sharp-heron-93

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Reworks #11684 (convergence item F) under the operator ruling of 2026-09-19, verbatim:
"there is no CPU demand - the processes we run are single core - do NOT overthink it please".

What changes

gunbc.runner_slot_allocation

  • cores_per_slot is an authored POLICY of 1 (gunbc_runner_cores_per_slot_policy), citing the ruling. It is not derived from memory_max and not from a vendor price list.
  • The memory-per-core ratio path is DELETED, not left beside it: gunbc_runner_memory_per_core, runner_memory_per_core_of, RunnerMemoryPerCoreResolution (incl. the now-unreachable MemoryPerCoreUnresolved arm), RunnerCoresPerSlotResolution, and the survey/ubicloud/ci_runner.types imports. gunbc_runner_cores_per_slot() returns a plain HardwareThreadCount.
  • Consequently the CPU axis admits 128 (cited Ampere M128-30 cores / 1) and stops binding; host_cpu_admitted_width and gunbc_derived_runner_slots_per_host lose their unreachable unresolved arms and return Int. Committed widths become the memory admissions: srv1 15, srv3/srv4 17 (srv2 still refused).
  • F-R3: only the CPU-survey refusal path is deleted. The memory-side unknown/zero conflation in gunbc_runner_slots_per_host (srv2's unestablished session charge reads as a scalar 0) remains, with its existing owner and trigger (a typed RunnerWidthResolution carried to every consumer) restated on the function.
  • The reason: authority row carries the ruling and the re-derived widths at its head; the stale "CPU binds" annotations are corrected.
  • The cores × memory_per_core ≥ memory_max conjunct of gunbc_runner_slot_allocation_wall_holds is deleted (it was the deleted derivation read back as a relation); the compiler-admission specimen recorded on it is kept on the wall's annotation.
  • Source cores_per_slot from a workload reserve, not a VM price list #11684's ByteSize reserve carrier (review 68200) is gone with the byte reserve.

CPUQuota: kept unset/unlimited, and the unset is affirmative. cpu_quota WAS derived from cores_per_slot; at a policy of 1 it would have rendered CPUQuota=100% on every runner unit and fabric cell slice, throttling parallel link/codegen. So:

  • RunnerSlotCpuQuota = CpuQuotaResolved { threads } | SlotCpuQuotaUnbounded (the refusing arm had no producer once the survey went); gunbc_runner_slot_cpu_quota() is SlotCpuQuotaUnbounded. cpu_weight=100 unchanged.
  • extdeps.systemd.unit_file gains DirectiveReset (the empty assignment, which systemd.resource-control(5) documents as the unset for CPUQuota= — infinity is not a value that parser admits) and the mint slice_cpu_quota_unbounded() → CPUQuota=.
  • The reset is emitted on both paths — gunbc.host_converge runner_cpu_boundary_knobs renders the per-slot drop-in knob as CPUQuota= under the new RunnerCpuBoundary::NoCpuCeiling arm, and gunbc.fabric_cell_effect always emits six directives with CPUQuota= as the sixth. What this PR does NOT claim (F-R1/F-R2): the reset does not clear a previously installed finite quota by production convergence. Both existing-quota repair paths are EXPLICITLY BLOCKED in source with cause and trigger: (a) fabric — an existing boundary differing only by a finite quota reconciles as MemberChanged → CellChangeNotYetRealizable, so the drift is detected and refused, not repaired; trigger is gunbc.change_realization admitting a bounded resource-boundary change at that address (set-property on the existing slice, never delete/recreate). (b) runner unit — the knob is the desired representation only; host_axis_caps does not enrol the CPU drop-in, host_runner_memory_provision has no live apply, and runner_unit_live_read yields no verdict for CPUQuota; trigger is enrolling the CPU drop-in in the actuated managed set plus a CPUQuotaPerSecUSec readback verdict (write CPUQuota=, read infinity, unreadable does not satisfy). The reported srv4 runner-unit census found unlimited quotas; other runner units and the fabric boundary (fabric-cell-srv3-06.slice) require their own readback; this PR does not establish fleet-wide quota conformity.
  • gunbc.ci_runner_placement runner_cpu_boundary_for now takes the slot's cpu_quota rather than a thread count, so the deploy boundary follows the row instead of re-deriving from cores_per_slot. RunnerCoreRatioUnresolved (no producer) deleted.
  • gunbc.fabric_cell_converge desired digest spells unbounded through the same systemd_cpu_quota_per_sec_identity(CpuQuotaUnbounded) the observed side uses, so a host printing infinity reconciles as satisfied by construction (and a host still at a finite quota reads as drift — visible, and refused rather than repaired per F-R1). fabric_cell_address_standing_for_quota and the SliceDirectivesQuotaSurveyRefused / CellCpuQuotaSurveyRefused arms are deleted with their producer.
  • gunbc.runner_microvm cell entitlement maps the unbounded row to CpuEntitlementRelativeOnly with a cause naming the ruling; the declared guest size is therefore the served answer.

Witnesses migrated to the policy, not the old literal: w_live_ratio_resolves_to_the_vendors_decimal_gigabytes and its two refusal siblings are deleted with their subject and replaced by w_cpu_axis_admits_the_catalog_core_count_at_one_core_per_slot; runner_slot_allocation_committed_widths_are_the_minimum_over_axes now pins 15/17/17 == memory admission with CPU at 128; new rows drive both quota arms by fixture (witness_boundary_add_renders_the_quota_line_under_both_arms, witness_an_unbounded_host_quota_is_satisfied_under_the_unbounded_row, the_cpu_quota_unset_renders_the_empty_assignment, the_cell_entitlement_is_the_slot_quota_arm_for_arm); the sudoers witness asserts CPUQuota= and no line-terminal % grant (narrowed per review 68479).

Sudoers projections regenerated through generated_artifact_gate main_wet.

Evidence

Scoped claim_batch runs on every touched witness module (all green; the one microvm row that flipped is recorded on the row). Pre-existing reds NOT introduced here: five rows in test.claim.host_allocation_conservation (srv4_commits_memory_width_disk_gated_at_apply, session_hosts_refuse_the_managed_width, derived_width_reproduces_the_surviving_host_classes, the_control_plane_reaches_the_commitment_only_where_memory_still_binds, the_uncharged_widths_are_what_the_charged_ones_are_measured_against) pin 16 GiB-era widths (29/25/21) and FAIL identically on a clean main ancestor (7261964); the module is outside required_gate_prefixes. Left as found — a separate repair.

Supersedes #11684.

🤖 Generated with Claude Code

…PUQuota stays unset

Operator ruling, 2026-09-19, verbatim: "there is no CPU demand - the
processes we run are single core - do NOT overthink it please".

gunbc.runner_slot_allocation: cores_per_slot is an authored POLICY of 1
(gunbc_runner_cores_per_slot_policy), not derived from memory_max and not
from a vendor price list. The memory-per-core ratio path is deleted with its
refusal vocabulary (RunnerMemoryPerCoreResolution incl. the unreachable
MemoryPerCoreUnresolved arm, RunnerCoresPerSlotResolution, the survey
imports). The CPU axis admits 128 and stops binding; committed widths become
the memory admissions: srv1 15, srv3/srv4 17. host_cpu_admitted_width and
the derived width lose their unreachable unresolved arms.

cpu_quota was derived from cores_per_slot and would have rendered
CPUQuota=100% -- a throttle on parallel link/codegen the ruling did not ask
for. RunnerSlotCpuQuota's refusing arm (no producer left) becomes
SlotCpuQuotaUnbounded, the production value. The unset is affirmative so a
previously installed finite quota is CLEARED, not merely no longer emitted:
extdeps.systemd.unit_file gains DirectiveReset (the empty assignment that
systemd.resource-control(5) documents as CPUQuota's unset) and
slice_cpu_quota_unbounded(); host_converge renders the per-slot drop-in knob
as `CPUQuota=` under RunnerCpuBoundary::NoCpuCeiling; fabric_cell_effect
always emits six directives with `CPUQuota=` as the sixth; the fabric cell
desired digest spells unbounded through extdeps.systemd's
systemd_cpu_quota_unbounded_identity so a host printing `infinity`
reconciles as satisfied. ci_runner_placement's boundary now follows the
slot's cpu_quota rather than re-deriving from cores_per_slot.
cpu_weight=100 unchanged.

Witnesses migrated to the policy: the vendor-ratio rows are deleted with
their subject; committed widths pin 15/17/17 == memory admission with CPU at
128; both quota arms are driven by fixture; the sudoers witness asserts
`CPUQuota=` and no `%` grant. Sudoers projections regenerated through
generated_artifact_gate main_wet.

Pre-existing reds not introduced here: five rows of
test.claim.host_allocation_conservation pin 16 GiB-era widths and FAIL
identically on a clean main ancestor (7261964); outside the required
gate.

Supersedes #11684.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gunbai-bot gunbai-bot Bot changed the title convergence F: cores_per_slot = 1 per operator ruling (CPU does not bind width) cores_per_slot = 1 per operator ruling; the CPU axis stops binding, CPUQuota stays unset Sep 19, 2026
@gunbai-bot
gunbai-bot Bot marked this pull request as ready for review September 19, 2026 15:02
gunbc-ci-auto-heal and others added 4 commits September 19, 2026 15:42
… blocked; narrow the % clause

- Merge origin/main so the heal job's dag/gunbc/heal_candidate.dag entry
  resolves on this branch.
- F-R1/F-R2: the CPUQuota= reset is a declared representation. Clearing a
  previously installed finite quota is EXPLICITLY BLOCKED on both paths,
  with cause and trigger named at gunbc.fabric_cell_effect (MemberChanged ->
  CellChangeNotYetRealizable; trigger: change_realization admitting a bounded
  resource-boundary change) and gunbc.host_converge (CPU drop-in not in the
  actuated managed set, no CPUQuotaPerSecUSec readback verdict; trigger:
  enrol + bind readback). No new apply machinery.
- F-R3: only the CPU-survey refusal path is deleted; the memory-side
  unknown/zero conflation in gunbc_runner_slots_per_host keeps its owner and
  RunnerWidthResolution trigger. Reason row: 'no finite ceiling is declared;
  the reset is emitted'.
- review 68479: the sudoers negative is a line-terminal %.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The srv4 runner-unit census is evidence about srv4's runner units; other
units and the fabric boundary need their own readback, and this PR does
not establish fleet-wide quota conformity.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The five retirement rows drove srv4-13 / srv1-13 as the first surplus slot
while the fleet committed 12. At 17/15 those are members, so the rows were
exercising membership, not retirement, and went red on the required floor.
The obligation itself is derived (runner_width_transition:
committed + 1 .. prior_width), so only the fixtures move: srv4-18/19,
srv1-16. 36/36 PASS scoped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Sep 19, 2026
Merged via the queue into main with commit 6a3c9d5 Sep 19, 2026
4 checks passed
@gunbai-bot
gunbai-bot Bot deleted the session/sharp-heron-93 branch September 19, 2026 22:43
@briansrls
briansrls restored the session/sharp-heron-93 branch September 19, 2026 22:46
briansrls pushed a commit that referenced this pull request Sep 19, 2026
The 2026-09-19 no-ceiling ruling (#11716) makes the production entitlement
RelativeOnly; the shape now admits the Work's threads unbounded on that arm
instead of refusing, since a relative share can refuse no count. The
guest-size derivation stall stays retired; main's edit to it is superseded by
runner_microvm_floor_fit_stall. Fleet-cost sentences re-pointed at the memory
axis, which binds now.
gunbai-bot Bot pushed a commit that referenced this pull request Sep 19, 2026
…rom a width

#11716 moved srv3/srv4's committed width to 17, so the width transition derives
retirees 18..21 while srv4_transition_interruption_authorization still enumerates
13..21. The data is right and stays right: the scope is the units the operator
named on 2026-09-12, and operator_reason is a record of that decision rather
than a live derivation -- neither is edited.

What was stale is the annotation above it, which said "srv4's committed width
becomes 12, so slots 13..21 are the surplus population". That derived the scope
from a width, so it rotted the moment the width moved, and a reader checking it
against the current allocation would find the premise false and the row correct.

The note now states the breadth and why it is the safe direction: an
authorization wider than the obligation interrupts nothing by itself, because
the obligation decides WHAT is retired and this row only decides what MAY be
interrupted when it is; a NARROWER authorization is the arm that refuses. It
also makes the enumeration's real argument load-bearing -- a range would have
silently RE-ACQUIRED 13..17 as the width moved back up, granting permission over
slots that are once again committed. An enumeration cannot change meaning
underneath the decision that authored it.

The per-subject coverage discriminator moves from srv4-13 to srv4-18. Both are
inside the authorization and either carries the claim; 18 is also a unit the
obligation now names, so a reader cannot mistake a claim about COVERAGE for a
claim about membership. Re-verified: stamping one availability across the
subjects still reddens exactly the two coverage claims at the new units.

Merged origin/main. Suites: host_retirement 75, capture 44, closure 35, route
13, legacy 9, provision 37, capacity_plan 24. Regen is a fixed point.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
gunbai-bot Bot pushed a commit that referenced this pull request Sep 20, 2026
…uota

cores_per_slot is HardwareThreadCount (policy of 1, #11716); the ephemeral unit emits no
CPUQuota word under SlotCpuQuotaUnbounded; the offer's thread axis is the quota when bounded and
the host's observed threads when not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants