Skip to content

Runner-slot caps: operator per-axis targets (16GiB x 5 guaranteed, swap 32GiB applied); de-fork swap knob onto the authority - #6463

Merged
briansrls merged 4 commits into
mainfrom
session/loyal-wren-398-slot-caps
Jul 11, 2026
Merged

briansrls merged 4 commits into
mainfrom
session/loyal-wren-398-slot-caps

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Jul 10, 2026 •

Copy link
Copy Markdown
Contributor

Runner-slot caps: operator per-axis targets applied, swap knob de-forked onto the authority row

Resolves the declared-vs-live cap fork found off the #6453 calibration point (merry-owl-649's read, escalated to the operator; ruling 2026-07-10): runner_slot_allocation.dag declared 32GiB MemorySwapMax + 8GiB MemoryMax, but the generated fleet-converge.sh applied MemorySwapMax=0 on every host, and the killing cgroup ran live at 24GiB — the authority row lying in two directions at once.

Per-axis targets (operator-set — explicitly NOT reconciled toward live)

  • MemorySwapMax = 32GiB (34359738368), applied: the row stays 32; the converge emit now derives from the row instead of a hardcoded "0" (the §3 fork: host_converge.dag's knob carried its own value beside the authority). Reconciling toward live-0 would have cemented the OOM.
  • MemoryMax = 16GiB (17179869184) per slot: fits the floor — width-1 needs ~14.7GiB (executor base + one shard) < 16GiB anon, with 32GiB swap catching brief spill. The prior 8GiB was below the floor's own ~8.3GiB base.
  • slots_per_host = 5: Σ = 5 × 16GiB = 80GiB = runner_slice_cap exactly — guaranteed mode, no oversubscription (the ≤ check holds with equality; the RED control at 6 slots still discriminates: 96 > 80).

Mechanism (one authority, three representations equal by construction)

  • RunnerHostDeployment gains per_runner_memory_swap_cap, fed from gunbc_runner_slot_desired().memory_swap_max — the swap knob now flows authority → deployment → emit, parallel to how MemoryMax already flowed. declared_runner_count() derives 80/16 = 5, so the count agrees with slots_per_host by construction at these values.
  • Verify-effective leg (the load-bearing part — the 32-vs-0 state existed because the apply was assumed): the emitted converge_per_slot_cap already does apply → systemctl show readback → verdict per unit; an authored-but-inert cap reads drifted, never green. This PR points that existing gate at the right expected values.
  • The live-read witness now carries the discriminating pair: the recorded pre-converge live read (swap=0, probed 2026-07-01) drifts against the desired target, and a target-matching read converges — the drift the fleet will show until the operator runs converge, detected as drift rather than blessed as converged-to-zero.
  • MemoryHigh = 15GiB, emitted and applied (operator follow-up ruling 2026-07-11): per_runner_memory_high_cap flows authority → deployment → knob (40-fleet-high.conf) with its own verify-effective pair (recorded live read high=infinity drifts against the 15GiB target; matching read converges). The authored max−1GiB relation, now a real cap: a hit at 15 throttles and spills to swap before the 16 kill line. Σ-swap (160GiB/host) is resolved by the operator's host-side provisioning guarantee; the Σ-swap ≤ host_swap_budget wall stays roadmap host-admission.

Receipts

  • 14/14 (+10/10 after the MemoryHigh fold-in) touched witnesses green by execution: allocation wall, boundary-exact, RED oversubscription control, emit pins (all three hosts at the targets, caps-before-widen ordering preserved), converge-delta keystone, live-read drift/converge pair, conservation suite, declared_runner_count_is_five, placement aggregate.
  • fleet-converge.sh regenerated via main_wet; the drift witness (committed == expected) is green, so the generated artifact matches the new authority.
  • Count pins consciously moved per the counting discipline: declared_runner_count_is_ten → is_five (2 sites); derives_ten renamed to the value-agnostic derives_declared_count (it compared symbolically already).

Post-merge sequence

  1. Operator runs the regenerated fleet-converge.sh (ctrl applies host-side — never from this repo).
  2. Readback receipts land 16GiB/32GiB live; gunbc_ci_runner_slot_memory_max_live (24GiB, 2026-07-05 receipt) updates on that real converge receipt, firing its own dissolve-on (live == declared → collapse to the single authority).
  3. The 8GiB × 10 shape returns as the after-S2b-base-shrink target, a future ruling.

Related: the rust-tests-from-CI drop (reclaiming the 10→5 worker cut) is merry-owl-649's lane, deliberately not folded in here.

briansrls and others added 2 commits July 10, 2026 23:56
…ed, swap 32GiB applied) and de-fork the swap knob onto the authority row

Operator ruling 2026-07-10 (via merry-owl-649), resolving the declared-vs-live
fork on runner_slot_allocation: MemorySwapMax target 32GiB APPLIED (the emitted
converge hardcoded MemorySwapMax=0 against the row's 32GiB — the knob now
derives from the row through RunnerHostDeployment.per_runner_memory_swap_cap);
MemoryMax target 16GiB/slot; slots_per_host 5. Sigma = 5x16 = 80GiB =
runner_slice_cap exactly — guaranteed mode, no oversubscription; width-1 floor
(~14.7GiB) fits the slot anon. declared_runner_count derives 80/16 = 5, so the
count stays coherent with the row by construction. The prior 8GiB x 10 shape is
the after-S2b target, not current.

row == converge emit == expected readback, and the emitted converge_per_slot_cap
already readback-verifies each apply (drifted on inert apply, the verify-effective
leg). The live-read witness now carries the discriminating pair: the recorded
pre-converge live read (swap=0) DRIFTS against the desired target; a
target-matching read converges. Count pins consciously moved: is_ten -> is_five.

14/14 touched witnesses green by execution (wall, boundary-exact, RED
oversubscription control, emit pins, delta keystone, live-read pair,
conservation, placement). memory_high preserves the authored max-1GiB relation
(15GiB; converge does not yet emit a MemoryHigh knob — row coherence only).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@gunbai-bot gunbai-bot Bot changed the title affected set processing Runner-slot caps: operator per-axis targets (16GiB x 5 guaranteed, swap 32GiB applied); de-fork swap knob onto the authority Jul 11, 2026
@gunbai-bot
gunbai-bot Bot marked this pull request as ready for review July 11, 2026 00:03
briansrls and others added 2 commits July 11, 2026 00:10
… authority -> deployment -> knob -> verify trio as swap

per_runner_memory_high_cap on RunnerHostDeployment fed from
gunbc_runner_slot_desired().memory_high (15GiB, the authored max-1GiB
relation); MemoryHigh legacy_converge_knob (40-fleet-high.conf) emitted per
slot on all three hosts; wall pins the value; the row-coherence-only caveat
is dropped from the authority disposition. Verify-effective pair added: the
recorded live read (high=infinity) DRIFTS against the 15GiB target, a
matching read converges. Sigma-swap (160GiB/host) resolved by operator
provisioning guarantee host-side; the Sigma-swap wall stays roadmap
host-admission. 10/10 touched witnesses green by execution; fleet-converge.sh
regenerated, drift gate green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@briansrls
briansrls merged commit e50a16f into main Jul 11, 2026
2 of 3 checks passed
@briansrls
briansrls deleted the session/loyal-wren-398-slot-caps branch July 11, 2026 00:15
briansrls added a commit that referenced this pull request Jul 11, 2026
…flushed raw markers)

All six conflicted files take origin/main verbatim: HEAD side was stale
pre-#6453/#6463/#6466 session WIP; main carries the merged truth.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant