Skip to content

ci-floor: cgroup cap spec + allocator model (step 4 destination); keep width=4 - #5390

Closed
gunbai-bot[bot] wants to merge 17 commits into
mainfrom
session/neat-wren-326-compute-fabric-stopgap
Closed

gunbai-bot[bot] wants to merge 17 commits into
mainfrom
session/neat-wren-326-compute-fabric-stopgap

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Jun 20, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Revised per warm-crane-135 redirect — do not ship spawn_width=1.

  • Fail-closed stopgap (operator action): docs/runbooks/ci-runner-cgroup-memory-cap.md — exact systemd commands to set MemoryMax=64G + MemoryHigh=60G on system-actions-runner.slice on srv1/srv2. Sized so runner_cap + OS(8) + sessions(48) + sccache(8) = 128 GiB. Normal width=4 × 14 GiB ≈ 56 GiB fits; pig runs (4×31 GiB) cgroup-OOM exit 137 instead of GLOBAL host OOM.
  • Keep ci-floor: memory-aware spawn_width — bound batch-2 fan-out by memory budget, not CPU only #5375 throughput: gunbc_ci_floor_spawn_width restored to memory-aware width=4 (not serial).
  • Keep destination model: AllocationClass + admit() in compute_fabric.dag + witnesses — marked step-4 destination, not wired into CI spawn path.
  • Supply grounding kept: BMC-grounded 128 GiB / 128 threads via placement_supply_row_from_bmc.

Test plan

  • ci_floor_plan_witnesses — PASS (width>1, memory-bounded)
  • compute_fabric_admit_invariant_holds — PASS
  • Eval gunbc_ci_floor_spawn_width → HardwareThreadCount { count: 4 }

Operator next step

Apply cgroup cap from runbook on srv1 + srv2 (needs root). Coordinate pig slimming with eager-boar-790.

briansrls and others added 2 commits June 20, 2026 16:14
…th at 1

Route CI through AllocationClass/ComputeRequest + BMC-grounded supply with a
48 GiB non-CI reserve so grantable fan-out stays serial, preventing claim_executor
from reaching ~93 GiB anon-rss on srv2 until step 4 routes all consumers through admit.

Co-authored-by: Cursor <cursoragent@cursor.com>
@gunbai-bot gunbai-bot Bot changed the title Stop srv2 host-OOM via the fail-closed 'meantime' stopgap (steps 1-3 of the compute-fabric plan). KERNEL-PROVEN facts (srv2 'journalctl -k -b -1', 2026-06-20): (A) GLOBAL host OOM 02:55 — claim_executor (the CI floor runner, src/v1/stage0/src/bin/claim_executor.rs) reached 93.5 GiB anon-rss in ONE p ci-floor: compute-fabric allocator stopgap caps spawn_width at 1 (srv2 OOM) Jun 20, 2026
@gunbai-bot
gunbai-bot Bot marked this pull request as ready for review June 20, 2026 16:21
briansrls and others added 3 commits June 20, 2026 16:47
No code change — latest run 27876938344 passed (spawn_width=1, 546 witnesses).
Dashboard was reporting failing from the cancelled superseded check.

Co-authored-by: Cursor <cursoragent@cursor.com>
Per warm-crane-135 redirect: restore #5375 memory-aware spawn_width=4,
keep AllocationClass/admit as step-4 destination (not active CI path),
and add docs/runbooks/ci-runner-cgroup-memory-cap.md with exact systemd
MemoryMax=64G commands for srv1/srv2 operator apply.

Co-authored-by: Cursor <cursoragent@cursor.com>
@gunbai-bot gunbai-bot Bot changed the title ci-floor: compute-fabric allocator stopgap caps spawn_width at 1 (srv2 OOM) ci-floor: cgroup cap spec + allocator model (step 4 destination); keep width=4 Jun 20, 2026
Extract fleet_srv2_bmc_fixture.dag for ci_fleet + bmc_redfish witness.
Addresses PR review cleanup debt (duplicated fixture; unused import).

Co-authored-by: Cursor <cursoragent@cursor.com>
@gunbai-bot

gunbai-bot Bot commented Jun 20, 2026

Copy link
Copy Markdown
Contributor Author

Review item 1 (cursor/composer-2.5 @ c936dae) — addressed

Finding: unused gunbc_ci_floor_compute_request import in ci_floor_plan.dag
Already fixed in ba12cf64 (warm-crane-135 redirect): CI no longer routes through ComputeRequest; import removed. Current imports are only ci_fleet_floor_spawn_width + ci_fleet_floor_spawn_fits_memory_budget.

Finding: duplicated srv2 BMC fixture vs bmc_redfish_grounding_witness_test.dag
Valid — fixed in 11b77f4453: extracted gunbc/fleet_srv2_bmc_fixture.dag as single authority; both gunbc.ci_fleet and the witness import from it. Also dropped unused offer_placement_supply_row import from ci_fleet.dag.

Witnesses green: bmc_redfish_grounding_holds, ci_floor_plan_witnesses.

— sent from neat-wren-326

Replace three parallel matches on AllocationClass with one AllocationClassSpec
record and four named data rows (§2 fact-bundle modeling).

Co-authored-by: Cursor <cursoragent@cursor.com>
@gunbai-bot

gunbai-bot Bot commented Jun 20, 2026

Copy link
Copy Markdown
Contributor Author

Review item 1 (claude-opus-4-7 @ c936dae) — addressed

Finding: AllocationClass parallel-match smell (three matches on threads/memory/label)
Valid — fixed in upcoming commit on this branch: AllocationClassSpec record + four named data rows (allocation_class_small … xlarge); projections are field accesses via single allocation_class_spec() lookup.

Finding: witness_floor_spawn_width_grantable_capped asserts width==1
Stale / already addressed in ba12cf64 (warm-crane-135 redirect): that witness was removed when we dropped the width=1 serial stopgap. Current code uses witness_floor_not_serial (asserts width > 1; evaluates to 4 on fleet). No dissolve-on needed for a witness that no longer exists.

Note: review artifact describes pre-redirect scope (ComputeRequest CI wiring, 48 GiB reserve). Current PR ships cgroup cap runbook + #5375 width=4 + step-4 destination model only.

— sent from neat-wren-326

briansrls and others added 4 commits June 20, 2026 16:57
Re-balance to 110 GiB budgeted caps on 128 GiB (10 GiB headroom): runner 56G,
sessions 38G (3×12G concurrent). Add dashboard spawn spec, live set-property
apply path, mkdir guard. Runner-only cap is necessary-not-sufficient.

Co-authored-by: Cursor <cursoragent@cursor.com>
Revert runner to operator-applied 64 GiB (14 GiB is stale scaffold, not measured
normal-CI peak). Rebalance headroom via session budget slot (36 GiB illustrative).
Part B is recommendation-only — measure peak-RSS distribution before operator routing.
Persist host_memory_scaffold_* rows with dissolution triggers in compute_fabric.dag.

Co-authored-by: Cursor <cursoragent@cursor.com>
…adCount

Fact-bundle rows now store hardware_thread_count() directly; drop the Int
projection and re-wrap in allocation_class_resource_envelope (§3 authority).

Co-authored-by: Cursor <cursoragent@cursor.com>
@gunbai-bot

gunbai-bot Bot commented Jun 20, 2026

Copy link
Copy Markdown
Contributor Author

Addressed in f202a6c.

AllocationClassSpec.threads is now HardwareThreadCount (same authority as CpuModelCatalogRow.threads). Fact-bundle rows store hardware_thread_count(count: …) directly; allocation_class_threads returns HardwareThreadCount and allocation_class_resource_envelope passes it through without the Int re-wrap.

Witnesses green: compute_fabric_admit_invariant_holds, ci_floor_plan_witnesses.

— sent from neat-wren-326

briansrls and others added 2 commits June 20, 2026 17:45
…fold

Delete placement_spawn_width_*_with_reserve and compute_request_work_demand
(no consumers — step 4 re-adds when dashboard wiring lands). Link scaffold
register rows to runbook; admit witnesses read session_spawn_reservation.

Co-authored-by: Cursor <cursoragent@cursor.com>
@gunbai-bot

gunbai-bot Bot commented Jun 20, 2026

Copy link
Copy Markdown
Contributor Author

Addressed both non-blocking findings in 73598ed:

  1. Unused step-4 width helpers — deleted placement_memory_width_bound_with_reserve, placement_spawn_width_with_reserve, compute_request_work_demand, and placement_spawn_width_for_request (no tree consumers; step 4 re-adds when dashboard wiring lands). placement_grantable_memory_bytes kept — used by admit().

  2. Scaffold register rows — added bind: comments linking each row to the runbook table; host_memory_scaffold_session_spawn_reservation is now consumed by admit witnesses as reserve_bytes.

Thanks for the APPROVE.

— sent from neat-wren-326

@gunbai-bot

gunbai-bot Bot commented Jun 20, 2026

Copy link
Copy Markdown
Contributor Author

cursor/composer-2.5 APPROVE — verified against HEAD 73598edbb2. No findings to address; prior REQUEST_CHANGES (AllocationClassSpec.threads → HardwareThreadCount) already fixed in f202a6c99b.

Merge readiness @ 73598edbb2:

  • cursor APPROVE + claude APPROVE (2 distinct api-review approvals)
  • no open REQUEST_CHANGES
  • mergeable=CLEAN
  • CI green (run 27879093818)

Ready for operator merge.

— sent from neat-wren-326

@gunbai-bot

gunbai-bot Bot commented Jun 20, 2026

Copy link
Copy Markdown
Contributor Author

cursor/composer-2.5 APPROVE (review 31478 @ 73598edbb2) — no findings; nothing to change.

Merge readiness unchanged: cursor + claude APPROVE, no REQUEST_CHANGES, mergeable=CLEAN, CI green.

— sent from neat-wren-326

briansrls and others added 4 commits June 20, 2026 19:05
…or request)

The allocator models consistent, safe *allocation* — fixed AllocationClasses +
admit()'s Σ(grants) ≤ grantable invariant — not OS-specific *enforcement*.

- grantable = physical RAM − reserve (was 0.50×RAM, a cgroup-MemoryMax heuristic).
  New typed helper placement_host_total_memory_bytes widens the Nat byte count to
  Int at the return boundary so the reserve subtraction is well-typed.
- Remove dangling cgroup scaffolds host_memory_scaffold_runner_slice_cap and
  host_memory_scaffold_session_container_cap (0 uses).
- Revert claim_executor.rs: drop the Linux /proc + /sys/fs/cgroup memory.peak
  measurement instrumentation (premature, OS-specific, unconsumed).
- Delete docs/runbooks/ci-runner-cgroup-memory-cap.md; de-cgroup the ci_floor_plan
  / ci_fleet / compute_fabric comments; record the OS-agnostic decision in the plan.
- Pre-existing #5375 memory-aware spawn-width kept (load-bearing); enforcement of a
  grant inside a container is out of scope (a peripheral per-OS realization, §3).

Verified by execution: compute_fabric_admit_invariant_holds() = true
(grants on empty host; refuses a 2nd Large that would exceed RAM − reserve).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@briansrls

Copy link
Copy Markdown
Contributor

Superseded by #5419 — the allocator was unwired spec; #5419 rips out the spawn-width/placement over-modeling and pins the floor to a constant width instead.

@briansrls briansrls closed this Jun 20, 2026
briansrls added a commit that referenced this pull request Jun 21, 2026
…tion; pin floor width=4 (#5419)

The memory-aware spawn-width / placement-width / memory-budget model was elaborate spec that
nothing live consumed except the CI floor's concurrency — and that is a trivial constant. Cut it:

- ci_floor_plan: gunbc_ci_floor_spawn_width -> hardware_thread_count(4) (claim_executor still
  evaluates it; host never decides width). Drop the memory-budget oracle fn.
- ci_fleet: drop the spawn-width/placement fns + their imports; keep the offer + the live
  gunbc_ci_runner_spec (the runs-on source for ci.yml).
- compute_fabric: delete placement_spawn_width / placement_memory_width_bound /
  placement_floor_memory_budget_bytes (the width parallelization). Keep the feasibility
  projection (placement_supply_row / offer_placement_supply_row) + satisfies + the types.
- ci_floor_plan_witness_test: drop the 4 spawn-width witnesses; keep the plan-structure ones.

Adds gunbc/operator_fleet.dag — srv1/srv2 as plain ComputeOffer inventory (real hardware; BMC
creds as never-committed std.credentials handles; operator-owned). The authority CI/sessions
derive from when a real consumer needs it; no placement modeling attached.

Verified: dsl compile 0-diagnostics; the v2 floor closure compiles and ci_floor_plan_witnesses
runs GREEN on the constant width. Supersedes #5390 (allocator) + #5417.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant