Skip to content

ci-floor: memory-aware spawn_width — bound batch-2 fan-out by memory budget, not CPU only - #5375

Merged
briansrls merged 10 commits into
mainfrom
session/warm-crane-135
Jun 20, 2026
Merged

briansrls merged 10 commits into
mainfrom
session/warm-crane-135

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Jun 20, 2026 •

Copy link
Copy Markdown
Contributor

What

placement_spawn_width was memory-blind: min(shard_count, hardware_threads). After #5356 took the floor to width=7, batch-2's concurrent corpus resolves (~14 GiB peak each) OOM-kill (exit 137) deterministically as the corpus grows — confirmed by eager-boar-790 (3 runs, both hosts) and implicated in srv2's memory-exhaustion reboots today.

Wire the already-modeled memory terms into the width — no new vocabulary:

spawn_width = min(shard_count, hardware_threads, memory_budget / per_unit_peak)
  • supply side: PlacementSupplyRow.ram_bytes (corrected to the measured 125 GiB; the 256 GiB stub made the bound inert)
  • demand side: ResourceEnvelope.memory = 14 GiB/shard, grounded in the OOM evidence
  • budget: conservative 0.50×RAM (the tightest host's cgroup MemoryMax), grounds-tighter when ctrl/BMC lands the real per-host cap

This is the connected model DESIGN §6 requires — CPU and memory are ONE width relationship, not a CPU bound patched after an OOM (§1 safety axis made structural).

Verified by execution

  • gunbc_ci_floor_spawn_width → HardwareThreadCount { count: 4 } (was 7; 4×14 GiB = 56 GiB ≤ 62.5 GiB budget, safe under srv1's ~65 GiB cap)
  • gunbc_ci_floor_spawn_fits_memory_budget → true (the §5 OOM oracle)
  • ci_floor_plan_witnesses → true — adds witness_floor_width_fits_memory_budget + witness_floor_width_memory_bounded, discriminating against the memory-blind regression

🤖 Generated with Claude Code

briansrls and others added 9 commits June 20, 2026 00:52
…minating witness

Floor ran serially because the corpus WorkDemand hardcoded shard_count=1, so
spawn_width = min(1, threads) = 1. Now the floor plan derives breadth from the
single-authority spec (one unit per gate) and the fleet supplies the thread cap:
spawn_width = min(breadth, hardware_threads). Verified by execution: 1 -> 7.

Witnesses: spawn_width tracks breadth (the relationship), and floor-not-serial
(discriminating — reverting to a constant shard_count goes RED).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
claude-opus-4-7 review on #5356: the +1 fold is just length(). Verified
length resolves on gunbc_ci_spec.gates across the v2->dsl bridge and still
yields 7; witness stays green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…roup budget, not CPU only

placement_spawn_width was memory-blind (min of shard_count and hardware_threads).
At width=7 the batch-2 corpus resolves OOM-killed (exit 137) deterministically as the
corpus grows (eager-boar-790, 3 runs). Wire the already-modeled memory terms
(PlacementSupplyRow.ram_bytes + ResourceEnvelope.memory) into the width:
spawn_width = min(shard_count, hardware_threads, memory_budget / per_unit_peak).

Verified by execution: floor spawn_width 7 -> 4 (4x14GiB=56GiB <= 0.50x125GiB budget,
safe under srv1's ~65GiB cgroup cap); gunbc_ci_floor_spawn_fits_memory_budget -> true;
ci_floor_plan_witnesses -> true (adds witness_floor_width_fits_memory_budget +
witness_floor_width_memory_bounded, discriminating vs the memory-blind regression).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… HEAD; main side == #5356 subset)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@briansrls
briansrls marked this pull request as ready for review June 20, 2026 03:46
@briansrls briansrls changed the title CI is extremely flakey ci-floor: memory-aware spawn_width — bound batch-2 fan-out by memory budget, not CPU only Jun 20, 2026
@briansrls
briansrls merged commit b3e67b0 into main Jun 20, 2026
1 of 2 checks passed
@briansrls
briansrls deleted the session/warm-crane-135 branch June 20, 2026 04:04
briansrls added a commit that referenced this pull request Jun 20, 2026
Per warm-crane-135 redirect: restore #5375 memory-aware spawn_width=4,
keep AllocationClass/admit as step-4 destination (not active CI path),
and add docs/runbooks/ci-runner-cgroup-memory-cap.md with exact systemd
MemoryMax=64G commands for srv1/srv2 operator apply.

Co-authored-by: Cursor <cursoragent@cursor.com>
briansrls added a commit that referenced this pull request Jun 20, 2026
…or request)

The allocator models consistent, safe *allocation* — fixed AllocationClasses +
admit()'s Σ(grants) ≤ grantable invariant — not OS-specific *enforcement*.

- grantable = physical RAM − reserve (was 0.50×RAM, a cgroup-MemoryMax heuristic).
  New typed helper placement_host_total_memory_bytes widens the Nat byte count to
  Int at the return boundary so the reserve subtraction is well-typed.
- Remove dangling cgroup scaffolds host_memory_scaffold_runner_slice_cap and
  host_memory_scaffold_session_container_cap (0 uses).
- Revert claim_executor.rs: drop the Linux /proc + /sys/fs/cgroup memory.peak
  measurement instrumentation (premature, OS-specific, unconsumed).
- Delete docs/runbooks/ci-runner-cgroup-memory-cap.md; de-cgroup the ci_floor_plan
  / ci_fleet / compute_fabric comments; record the OS-agnostic decision in the plan.
- Pre-existing #5375 memory-aware spawn-width kept (load-bearing); enforcement of a
  grant inside a container is out of scope (a peripheral per-OS realization, §3).

Verified by execution: compute_fabric_admit_invariant_holds() = true
(grants on empty host; refuses a 2nd Large that would exceed RAM − reserve).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
briansrls added a commit that referenced this pull request Jun 21, 2026
…quick-ant)

The §3 invariant conflated two phases. spawn_width is the discovery-corpus
RUN-phase shard width (witnesses against the prebuilt binary — cores∧mem-bound,
pids-light), so it does NOT multiply per_build_pids; width-up is pids-safe on
its own. The pids crash is the BUILD phase: concurrent_builds × per_build_pids
≤ pids_cap, coupling host-packing × fan-out, not spawn_width. Two invariants,
not one product. Also: #5375 memory-aware was superseded by #5419's pin (#5444
re-adds the term).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
briansrls added a commit that referenced this pull request Jun 21, 2026
…r one authority) (#5462)

* docs/plans: compute-envelope-model — one authority for the CI fleet's resource dimensions

Plan doc resolving the §1 ROADMAP "CI on compute fabric" pointer. Models the
bimodal crash-or-idle pathology as one root (N hand-tuned resource dimensions,
no single ResourceEnvelope authority) and the §3 fix: derive every knob —
spawn_width, fan-out, TasksMax, jobserver, MemoryMax — from one measured
envelope, with the public(shape)/ctrl(realization) split. Co-owned warm-lark-306
+ quick-ant-298 (§1 lead); CC bright-stag-194 (ROADMAP + test profile).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* compute-envelope-model: correct current width = pinned 4 (quick-ant verified)

Memory-aware width model was removed as unwired (#5419); current floor width
is a pinned constant 4 (ci_floor_plan.dag:292), ~3% of 128c. #5444 re-adds the
memory term (not yet on main). The lever lifts both the pin and the shard_count
cap from the envelope. Fresh spawn-width slice to be reopen-scoped by quick-ant.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ROADMAP §1: point the CI-on-fabric chore-line at the compute-envelope plan doc

Makes #5462 self-contained (doc + its own pointer, atomic, no orphan window).
Different line from the §1 nightly-reframe edits — no collision.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* compute-envelope-model: split RUN-phase width from BUILD-phase pids (quick-ant)

The §3 invariant conflated two phases. spawn_width is the discovery-corpus
RUN-phase shard width (witnesses against the prebuilt binary — cores∧mem-bound,
pids-light), so it does NOT multiply per_build_pids; width-up is pids-safe on
its own. The pids crash is the BUILD phase: concurrent_builds × per_build_pids
≤ pids_cap, coupling host-packing × fan-out, not spawn_width. Two invariants,
not one product. Also: #5375 memory-aware was superseded by #5419's pin (#5444
re-adds the term).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ROADMAP §1: fix chore-line spawn_width conflation (RUN/BUILD split)

main's compact bullet (#5459) said 'spawn_width (memory- & pids-aware)' — the
same conflation quick-ant corrected in compute-envelope-model.md. spawn_width is
RUN-phase (cores∧mem, prebuilt binary, pids-light); pids binds the BUILD phase
(fan-out × concurrent-builds). Keeps ROADMAP consistent with the doc in one PR.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ROADMAP §1: adopt quick-ant's precise spawn_width de-conflation wording

Explicit RUN/BUILD phase labels — spawn_width = cores∧mem (RUN, prebuilt
binary, pids-light); pids burst belongs on per-build fan-out (BUILD,
concurrent_builds × per-build-pids ≤ TasksMax). No "pids" on the spawn_width
bullet at all (the conflation's verification check).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ROADMAP §1: shorten CI-on-fabric chore-line to a summary + doc pointer

Per operator: ROADMAP lines stay scannable progress summaries; density goes in
plan docs (its own stated rule — no restating detail). The 1175-char inline
chore-list/diagnosis/formulas now live in compute-envelope-model.md; the line
is a one-sentence summary + pointer. Density preserved in the doc, not lost.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Brian Searls <briansrls@gunb.ai>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant