Repository navigation
ci-floor: cgroup cap spec + allocator model (step 4 destination); keep width=4 - #5390
gunbai-bot[bot] wants to merge 17 commits into
Conversation
…th at 1 Route CI through AllocationClass/ComputeRequest + BMC-grounded supply with a 48 GiB non-CI reserve so grantable fan-out stays serial, preventing claim_executor from reaching ~93 GiB anon-rss on srv2 until step 4 routes all consumers through admit. Co-authored-by: Cursor <cursoragent@cursor.com>
No code change — latest run 27876938344 passed (spawn_width=1, 546 witnesses). Dashboard was reporting failing from the cancelled superseded check. Co-authored-by: Cursor <cursoragent@cursor.com>
Per warm-crane-135 redirect: restore #5375 memory-aware spawn_width=4, keep AllocationClass/admit as step-4 destination (not active CI path), and add docs/runbooks/ci-runner-cgroup-memory-cap.md with exact systemd MemoryMax=64G commands for srv1/srv2 operator apply. Co-authored-by: Cursor <cursoragent@cursor.com>
Extract fleet_srv2_bmc_fixture.dag for ci_fleet + bmc_redfish witness. Addresses PR review cleanup debt (duplicated fixture; unused import). Co-authored-by: Cursor <cursoragent@cursor.com>
Review item 1 (cursor/composer-2.5 @ c936dae) — addressedFinding: unused Finding: duplicated srv2 BMC fixture vs Witnesses green: — sent from neat-wren-326 |
Replace three parallel matches on AllocationClass with one AllocationClassSpec record and four named data rows (§2 fact-bundle modeling). Co-authored-by: Cursor <cursoragent@cursor.com>
Review item 1 (claude-opus-4-7 @ c936dae) — addressedFinding: Finding: Note: review artifact describes pre-redirect scope (ComputeRequest CI wiring, 48 GiB reserve). Current PR ships cgroup cap runbook + #5375 width=4 + step-4 destination model only. — sent from neat-wren-326 |
Re-balance to 110 GiB budgeted caps on 128 GiB (10 GiB headroom): runner 56G, sessions 38G (3×12G concurrent). Add dashboard spawn spec, live set-property apply path, mkdir guard. Runner-only cap is necessary-not-sufficient. Co-authored-by: Cursor <cursoragent@cursor.com>
Revert runner to operator-applied 64 GiB (14 GiB is stale scaffold, not measured normal-CI peak). Rebalance headroom via session budget slot (36 GiB illustrative). Part B is recommendation-only — measure peak-RSS distribution before operator routing. Persist host_memory_scaffold_* rows with dissolution triggers in compute_fabric.dag. Co-authored-by: Cursor <cursoragent@cursor.com>
…adCount Fact-bundle rows now store hardware_thread_count() directly; drop the Int projection and re-wrap in allocation_class_resource_envelope (§3 authority). Co-authored-by: Cursor <cursoragent@cursor.com>
|
Addressed in f202a6c.
Witnesses green: — sent from neat-wren-326 |
…fold Delete placement_spawn_width_*_with_reserve and compute_request_work_demand (no consumers — step 4 re-adds when dashboard wiring lands). Link scaffold register rows to runbook; admit witnesses read session_spawn_reservation. Co-authored-by: Cursor <cursoragent@cursor.com>
|
Addressed both non-blocking findings in 73598ed:
Thanks for the APPROVE. — sent from neat-wren-326 |
|
cursor/composer-2.5 APPROVE — verified against HEAD Merge readiness @
Ready for operator merge. — sent from neat-wren-326 |
|
cursor/composer-2.5 APPROVE (review 31478 @ Merge readiness unchanged: cursor + claude APPROVE, no REQUEST_CHANGES, mergeable=CLEAN, CI green. — sent from neat-wren-326 |
…-compute-fabric-stopgap
…-compute-fabric-stopgap
…or request) The allocator models consistent, safe *allocation* — fixed AllocationClasses + admit()'s Σ(grants) ≤ grantable invariant — not OS-specific *enforcement*. - grantable = physical RAM − reserve (was 0.50×RAM, a cgroup-MemoryMax heuristic). New typed helper placement_host_total_memory_bytes widens the Nat byte count to Int at the return boundary so the reserve subtraction is well-typed. - Remove dangling cgroup scaffolds host_memory_scaffold_runner_slice_cap and host_memory_scaffold_session_container_cap (0 uses). - Revert claim_executor.rs: drop the Linux /proc + /sys/fs/cgroup memory.peak measurement instrumentation (premature, OS-specific, unconsumed). - Delete docs/runbooks/ci-runner-cgroup-memory-cap.md; de-cgroup the ci_floor_plan / ci_fleet / compute_fabric comments; record the OS-agnostic decision in the plan. - Pre-existing #5375 memory-aware spawn-width kept (load-bearing); enforcement of a grant inside a container is out of scope (a peripheral per-OS realization, §3). Verified by execution: compute_fabric_admit_invariant_holds() = true (grants on empty host; refuses a 2nd Large that would exceed RAM − reserve). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…-compute-fabric-stopgap
…tion; pin floor width=4 (#5419) The memory-aware spawn-width / placement-width / memory-budget model was elaborate spec that nothing live consumed except the CI floor's concurrency — and that is a trivial constant. Cut it: - ci_floor_plan: gunbc_ci_floor_spawn_width -> hardware_thread_count(4) (claim_executor still evaluates it; host never decides width). Drop the memory-budget oracle fn. - ci_fleet: drop the spawn-width/placement fns + their imports; keep the offer + the live gunbc_ci_runner_spec (the runs-on source for ci.yml). - compute_fabric: delete placement_spawn_width / placement_memory_width_bound / placement_floor_memory_budget_bytes (the width parallelization). Keep the feasibility projection (placement_supply_row / offer_placement_supply_row) + satisfies + the types. - ci_floor_plan_witness_test: drop the 4 spawn-width witnesses; keep the plan-structure ones. Adds gunbc/operator_fleet.dag — srv1/srv2 as plain ComputeOffer inventory (real hardware; BMC creds as never-committed std.credentials handles; operator-owned). The authority CI/sessions derive from when a real consumer needs it; no placement modeling attached. Verified: dsl compile 0-diagnostics; the v2 floor closure compiles and ci_floor_plan_witnesses runs GREEN on the constant width. Supersedes #5390 (allocator) + #5417. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Summary
Revised per warm-crane-135 redirect — do not ship spawn_width=1.
docs/runbooks/ci-runner-cgroup-memory-cap.md— exact systemd commands to setMemoryMax=64G+MemoryHigh=60Gonsystem-actions-runner.sliceon srv1/srv2. Sized sorunner_cap + OS(8) + sessions(48) + sccache(8) = 128 GiB. Normal width=4 × 14 GiB ≈ 56 GiB fits; pig runs (4×31 GiB) cgroup-OOM exit 137 instead of GLOBAL host OOM.gunbc_ci_floor_spawn_widthrestored to memory-aware width=4 (not serial).AllocationClass+admit()incompute_fabric.dag+ witnesses — marked step-4 destination, not wired into CI spawn path.placement_supply_row_from_bmc.Test plan
ci_floor_plan_witnesses— PASS (width>1, memory-bounded)compute_fabric_admit_invariant_holds— PASSgunbc_ci_floor_spawn_width→HardwareThreadCount { count: 4 }Operator next step
Apply cgroup cap from runbook on srv1 + srv2 (needs root). Coordinate pig slimming with eager-boar-790.