Repository navigation
Conversation
…th at 1 Route CI through AllocationClass/ComputeRequest + BMC-grounded supply with a 48 GiB non-CI reserve so grantable fan-out stays serial, preventing claim_executor from reaching ~93 GiB anon-rss on srv2 until step 4 routes all consumers through admit. Co-authored-by: Cursor <cursoragent@cursor.com>
No code change — latest run 27876938344 passed (spawn_width=1, 546 witnesses). Dashboard was reporting failing from the cancelled superseded check. Co-authored-by: Cursor <cursoragent@cursor.com>
Per warm-crane-135 redirect: restore #5375 memory-aware spawn_width=4, keep AllocationClass/admit as step-4 destination (not active CI path), and add docs/runbooks/ci-runner-cgroup-memory-cap.md with exact systemd MemoryMax=64G commands for srv1/srv2 operator apply. Co-authored-by: Cursor <cursoragent@cursor.com>
Extract fleet_srv2_bmc_fixture.dag for ci_fleet + bmc_redfish witness. Addresses PR review cleanup debt (duplicated fixture; unused import). Co-authored-by: Cursor <cursoragent@cursor.com>
Replace three parallel matches on AllocationClass with one AllocationClassSpec record and four named data rows (§2 fact-bundle modeling). Co-authored-by: Cursor <cursoragent@cursor.com>
Re-balance to 110 GiB budgeted caps on 128 GiB (10 GiB headroom): runner 56G, sessions 38G (3×12G concurrent). Add dashboard spawn spec, live set-property apply path, mkdir guard. Runner-only cap is necessary-not-sufficient. Co-authored-by: Cursor <cursoragent@cursor.com>
Revert runner to operator-applied 64 GiB (14 GiB is stale scaffold, not measured normal-CI peak). Rebalance headroom via session budget slot (36 GiB illustrative). Part B is recommendation-only — measure peak-RSS distribution before operator routing. Persist host_memory_scaffold_* rows with dissolution triggers in compute_fabric.dag. Co-authored-by: Cursor <cursoragent@cursor.com>
…adCount Fact-bundle rows now store hardware_thread_count() directly; drop the Int projection and re-wrap in allocation_class_resource_envelope (§3 authority). Co-authored-by: Cursor <cursoragent@cursor.com>
…fold Delete placement_spawn_width_*_with_reserve and compute_request_work_demand (no consumers — step 4 re-adds when dashboard wiring lands). Link scaffold register rows to runbook; admit witnesses read session_spawn_reservation. Co-authored-by: Cursor <cursoragent@cursor.com>
…-compute-fabric-stopgap
…-compute-fabric-stopgap
…or request) The allocator models consistent, safe *allocation* — fixed AllocationClasses + admit()'s Σ(grants) ≤ grantable invariant — not OS-specific *enforcement*. - grantable = physical RAM − reserve (was 0.50×RAM, a cgroup-MemoryMax heuristic). New typed helper placement_host_total_memory_bytes widens the Nat byte count to Int at the return boundary so the reserve subtraction is well-typed. - Remove dangling cgroup scaffolds host_memory_scaffold_runner_slice_cap and host_memory_scaffold_session_container_cap (0 uses). - Revert claim_executor.rs: drop the Linux /proc + /sys/fs/cgroup memory.peak measurement instrumentation (premature, OS-specific, unconsumed). - Delete docs/runbooks/ci-runner-cgroup-memory-cap.md; de-cgroup the ci_floor_plan / ci_fleet / compute_fabric comments; record the OS-agnostic decision in the plan. - Pre-existing #5375 memory-aware spawn-width kept (load-bearing); enforcement of a grant inside a container is out of scope (a peripheral per-OS realization, §3). Verified by execution: compute_fabric_admit_invariant_holds() = true (grants on empty host; refuses a 2nd Large that would exceed RAM − reserve). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…-compute-fabric-stopgap
…ric authority (slice 2) The keystone the allocator + CI-as-placement both need: srv1 and srv2 modeled as real ComputeHost/ComputeOffer values (128c/128t Ampere Altra M128, 128 GiB DDR4, ASRock ALTRAD8UD-1L2T), grounded on live Redfish probes (srv2 reuses fleet_srv2_bmc_fixture; srv1 added). Per-host PlacementSupplyRow via placement_supply_row_from_bmc feeds the #5390 allocator's admit(). Replaces (next slice) the synthetic single gunbc_ci_fleet_offer. Public in gunbc — skipping the private ctrl repo. The privacy boundary is handled in-repo: rotated BMC credentials are std.credentials CredentialFlow.Stored HANDLES (secret name only, value never committed); operator-private distro/version left absent (resolved at runtime). The operator owns the fleet (provider = the operator; OwnedMarginalZero). Witnesses (operator_fleet_test.dag), green by execution: - srv1_placement_threads_grounded: 128 hw threads (real BMC), not the 64-thread stub. - srv1_admits_medium_on_empty_host: a Medium fits an empty srv1 (128 GiB − reserve). - srv1_refuses_when_reserve_exceeds_ram: the admit Σ≤grantable invariant on real RAM. Includes the umbrella plan docs/plans/operator-fleet-on-compute-fabric.md. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Contributor
Author
|
Superseded by #5419, which keeps operator_fleet as plain inventory and cuts the placement/allocator over-modeling (CI green on a constant floor width). |
briansrls
added a commit
that referenced
this pull request
Jun 21, 2026
…tion; pin floor width=4 (#5419) The memory-aware spawn-width / placement-width / memory-budget model was elaborate spec that nothing live consumed except the CI floor's concurrency — and that is a trivial constant. Cut it: - ci_floor_plan: gunbc_ci_floor_spawn_width -> hardware_thread_count(4) (claim_executor still evaluates it; host never decides width). Drop the memory-budget oracle fn. - ci_fleet: drop the spawn-width/placement fns + their imports; keep the offer + the live gunbc_ci_runner_spec (the runs-on source for ci.yml). - compute_fabric: delete placement_spawn_width / placement_memory_width_bound / placement_floor_memory_budget_bytes (the width parallelization). Keep the feasibility projection (placement_supply_row / offer_placement_supply_row) + satisfies + the types. - ci_floor_plan_witness_test: drop the 4 spawn-width witnesses; keep the plan-structure ones. Adds gunbc/operator_fleet.dag — srv1/srv2 as plain ComputeOffer inventory (real hardware; BMC creds as never-committed std.credentials handles; operator-owned). The authority CI/sessions derive from when a real consumer needs it; no placement modeling attached. Verified: dsl compile 0-diagnostics; the v2 floor closure compiles and ci_floor_plan_witnesses runs GREEN on the constant width. Supersedes #5390 (allocator) + #5417. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Carries forward the compute-fabric allocator work from #5390 (being closed) and adds slice 2 of
the fleet-on-fabric plan: srv1/srv2 as the real, BMC-grounded fabric authority. Sibling to the
access PR #5415 (disjoint files). Plan:
docs/plans/operator-fleet-on-compute-fabric.md.Allocator (from #5390 — the srv2-OOM fix)
AllocationClass(fixed Small/Medium/Large/XLarge) +admit(host, live, request)enforcing the onemissing invariant —
Σ(live allocations) + request ≤ grantable— so the fleet can no longerovercommit physical RAM (the 625 GiB-on-125 GiB livelock). OS-agnostic allocation (no cgroup); see
docs/plans/compute-fabric-allocator.md.Slice 2 —
dsl/gunbc/operator_fleet.dag(the keystone)srv1 + srv2 modeled as real
ComputeHost/ComputeOffer(128c/128t Ampere Altra M128, 128 GiB DDR4,ASRock ALTRAD8UD-1L2T), grounded on live Redfish probes (srv2 reuses
fleet_srv2_bmc_fixture; srv1added). Per-host
PlacementSupplyRow(placement_supply_row_from_bmc) feeds the allocator'sadmit()— replacing the synthetic singlegunbc_ci_fleet_offer.Privacy without ctrl (public-in-gunbc): rotated BMC credentials are
std.credentialsCredentialFlow.Storedhandles (secret name only, value never committed); operator-privatedistro/version left absent (resolved at runtime); the operator owns the fleet (
OwnedMarginalZero).Verification (DESIGN §5 — green by execution)
0 diagnostics.operator_fleet_test.dag, green by execution:srv1_placement_threads_grounded— 128 hw threads (real BMC), not the 64-thread stub; perturb→64 goes red.srv1_admits_medium_on_empty_host— a Medium fits an empty srv1.srv1_refuses_when_reserve_exceeds_ram— theadmitΣ≤grantable invariant on real RAM.Next (plan §5)
Slice 3: point CI's supply at the real per-host fleet (retire the synthetic offer). Slice 4:
CI-as-placement (per-host runner selection). Plus the access-policy layer (who may place/admin a host).
🤖 Generated with Claude Code