Repository navigation
ci-floor: memory-aware spawn_width — bound batch-2 fan-out by memory budget, not CPU only - #5375
Merged
Merged
Conversation
…minating witness Floor ran serially because the corpus WorkDemand hardcoded shard_count=1, so spawn_width = min(1, threads) = 1. Now the floor plan derives breadth from the single-authority spec (one unit per gate) and the fleet supplies the thread cap: spawn_width = min(breadth, hardware_threads). Verified by execution: 1 -> 7. Witnesses: spawn_width tracks breadth (the relationship), and floor-not-serial (discriminating — reverting to a constant shard_count goes RED). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
claude-opus-4-7 review on #5356: the +1 fold is just length(). Verified length resolves on gunbc_ci_spec.gates across the v2->dsl bridge and still yields 7; witness stays green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…roup budget, not CPU only placement_spawn_width was memory-blind (min of shard_count and hardware_threads). At width=7 the batch-2 corpus resolves OOM-killed (exit 137) deterministically as the corpus grows (eager-boar-790, 3 runs). Wire the already-modeled memory terms (PlacementSupplyRow.ram_bytes + ResourceEnvelope.memory) into the width: spawn_width = min(shard_count, hardware_threads, memory_budget / per_unit_peak). Verified by execution: floor spawn_width 7 -> 4 (4x14GiB=56GiB <= 0.50x125GiB budget, safe under srv1's ~65GiB cgroup cap); gunbc_ci_floor_spawn_fits_memory_budget -> true; ci_floor_plan_witnesses -> true (adds witness_floor_width_fits_memory_budget + witness_floor_width_memory_bounded, discriminating vs the memory-blind regression). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… HEAD; main side == #5356 subset) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
briansrls
marked this pull request as ready for review
June 20, 2026 03:46
briansrls
added a commit
that referenced
this pull request
Jun 20, 2026
Per warm-crane-135 redirect: restore #5375 memory-aware spawn_width=4, keep AllocationClass/admit as step-4 destination (not active CI path), and add docs/runbooks/ci-runner-cgroup-memory-cap.md with exact systemd MemoryMax=64G commands for srv1/srv2 operator apply. Co-authored-by: Cursor <cursoragent@cursor.com>
3 tasks done
briansrls
added a commit
that referenced
this pull request
Jun 20, 2026
…or request) The allocator models consistent, safe *allocation* — fixed AllocationClasses + admit()'s Σ(grants) ≤ grantable invariant — not OS-specific *enforcement*. - grantable = physical RAM − reserve (was 0.50×RAM, a cgroup-MemoryMax heuristic). New typed helper placement_host_total_memory_bytes widens the Nat byte count to Int at the return boundary so the reserve subtraction is well-typed. - Remove dangling cgroup scaffolds host_memory_scaffold_runner_slice_cap and host_memory_scaffold_session_container_cap (0 uses). - Revert claim_executor.rs: drop the Linux /proc + /sys/fs/cgroup memory.peak measurement instrumentation (premature, OS-specific, unconsumed). - Delete docs/runbooks/ci-runner-cgroup-memory-cap.md; de-cgroup the ci_floor_plan / ci_fleet / compute_fabric comments; record the OS-agnostic decision in the plan. - Pre-existing #5375 memory-aware spawn-width kept (load-bearing); enforcement of a grant inside a container is out of scope (a peripheral per-OS realization, §3). Verified by execution: compute_fabric_admit_invariant_holds() = true (grants on empty host; refuses a 2nd Large that would exceed RAM − reserve). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
briansrls
added a commit
that referenced
this pull request
Jun 21, 2026
…quick-ant) The §3 invariant conflated two phases. spawn_width is the discovery-corpus RUN-phase shard width (witnesses against the prebuilt binary — cores∧mem-bound, pids-light), so it does NOT multiply per_build_pids; width-up is pids-safe on its own. The pids crash is the BUILD phase: concurrent_builds × per_build_pids ≤ pids_cap, coupling host-packing × fan-out, not spawn_width. Two invariants, not one product. Also: #5375 memory-aware was superseded by #5419's pin (#5444 re-adds the term). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
briansrls
added a commit
that referenced
this pull request
Jun 21, 2026
…r one authority) (#5462) * docs/plans: compute-envelope-model — one authority for the CI fleet's resource dimensions Plan doc resolving the §1 ROADMAP "CI on compute fabric" pointer. Models the bimodal crash-or-idle pathology as one root (N hand-tuned resource dimensions, no single ResourceEnvelope authority) and the §3 fix: derive every knob — spawn_width, fan-out, TasksMax, jobserver, MemoryMax — from one measured envelope, with the public(shape)/ctrl(realization) split. Co-owned warm-lark-306 + quick-ant-298 (§1 lead); CC bright-stag-194 (ROADMAP + test profile). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * compute-envelope-model: correct current width = pinned 4 (quick-ant verified) Memory-aware width model was removed as unwired (#5419); current floor width is a pinned constant 4 (ci_floor_plan.dag:292), ~3% of 128c. #5444 re-adds the memory term (not yet on main). The lever lifts both the pin and the shard_count cap from the envelope. Fresh spawn-width slice to be reopen-scoped by quick-ant. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ROADMAP §1: point the CI-on-fabric chore-line at the compute-envelope plan doc Makes #5462 self-contained (doc + its own pointer, atomic, no orphan window). Different line from the §1 nightly-reframe edits — no collision. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * compute-envelope-model: split RUN-phase width from BUILD-phase pids (quick-ant) The §3 invariant conflated two phases. spawn_width is the discovery-corpus RUN-phase shard width (witnesses against the prebuilt binary — cores∧mem-bound, pids-light), so it does NOT multiply per_build_pids; width-up is pids-safe on its own. The pids crash is the BUILD phase: concurrent_builds × per_build_pids ≤ pids_cap, coupling host-packing × fan-out, not spawn_width. Two invariants, not one product. Also: #5375 memory-aware was superseded by #5419's pin (#5444 re-adds the term). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ROADMAP §1: fix chore-line spawn_width conflation (RUN/BUILD split) main's compact bullet (#5459) said 'spawn_width (memory- & pids-aware)' — the same conflation quick-ant corrected in compute-envelope-model.md. spawn_width is RUN-phase (cores∧mem, prebuilt binary, pids-light); pids binds the BUILD phase (fan-out × concurrent-builds). Keeps ROADMAP consistent with the doc in one PR. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ROADMAP §1: adopt quick-ant's precise spawn_width de-conflation wording Explicit RUN/BUILD phase labels — spawn_width = cores∧mem (RUN, prebuilt binary, pids-light); pids burst belongs on per-build fan-out (BUILD, concurrent_builds × per-build-pids ≤ TasksMax). No "pids" on the spawn_width bullet at all (the conflation's verification check). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ROADMAP §1: shorten CI-on-fabric chore-line to a summary + doc pointer Per operator: ROADMAP lines stay scannable progress summaries; density goes in plan docs (its own stated rule — no restating detail). The 1175-char inline chore-list/diagnosis/formulas now live in compute-envelope-model.md; the line is a one-sentence summary + pointer. Density preserved in the doc, not lost. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Brian Searls <briansrls@gunb.ai> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
placement_spawn_widthwas memory-blind:min(shard_count, hardware_threads). After #5356 took the floor to width=7, batch-2's concurrent corpus resolves (~14 GiB peak each) OOM-kill (exit 137) deterministically as the corpus grows — confirmed by eager-boar-790 (3 runs, both hosts) and implicated in srv2's memory-exhaustion reboots today.Wire the already-modeled memory terms into the width — no new vocabulary:
PlacementSupplyRow.ram_bytes(corrected to the measured 125 GiB; the 256 GiB stub made the bound inert)ResourceEnvelope.memory= 14 GiB/shard, grounded in the OOM evidenceThis is the connected model DESIGN §6 requires — CPU and memory are ONE width relationship, not a CPU bound patched after an OOM (§1 safety axis made structural).
Verified by execution
gunbc_ci_floor_spawn_width→HardwareThreadCount { count: 4 }(was 7; 4×14 GiB = 56 GiB ≤ 62.5 GiB budget, safe under srv1's ~65 GiB cap)gunbc_ci_floor_spawn_fits_memory_budget→true(the §5 OOM oracle)ci_floor_plan_witnesses→true— addswitness_floor_width_fits_memory_budget+witness_floor_width_memory_bounded, discriminating against the memory-blind regression🤖 Generated with Claude Code