Repository navigation
compute-fabric: cut the unwired spawn-width/placement over-modeling; pin floor width=4; keep fleet inventory - #5419
Merged
Merged
Conversation
…tion; pin floor width=4 The memory-aware spawn-width / placement-width / memory-budget model was elaborate spec that nothing live consumed except the CI floor's concurrency — and that is a trivial constant. Cut it: - ci_floor_plan: gunbc_ci_floor_spawn_width -> hardware_thread_count(4) (claim_executor still evaluates it; host never decides width). Drop the memory-budget oracle fn. - ci_fleet: drop the spawn-width/placement fns + their imports; keep the offer + the live gunbc_ci_runner_spec (the runs-on source for ci.yml). - compute_fabric: delete placement_spawn_width / placement_memory_width_bound / placement_floor_memory_budget_bytes (the width parallelization). Keep the feasibility projection (placement_supply_row / offer_placement_supply_row) + satisfies + the types. - ci_floor_plan_witness_test: drop the 4 spawn-width witnesses; keep the plan-structure ones. Adds gunbc/operator_fleet.dag — srv1/srv2 as plain ComputeOffer inventory (real hardware; BMC creds as never-committed std.credentials handles; operator-owned). The authority CI/sessions derive from when a real consumer needs it; no placement modeling attached. Verified: dsl compile 0-diagnostics; the v2 floor closure compiles and ci_floor_plan_witnesses runs GREEN on the constant width. Supersedes #5390 (allocator) + #5417. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This was referenced Jun 20, 2026
briansrls
added a commit
that referenced
this pull request
Jun 21, 2026
Adapt to #5419 (memory-aware spawn-width removed, floor width pinned to 4): drop the re-added placement/budget-units model; keep the run_walk host fix and re-express OOM safety as a flat topology policy — at most one heavy whole-tree resolve per readiness layer (source_root_ingest serialized behind the corpus). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
briansrls
added a commit
that referenced
this pull request
Jun 21, 2026
…s even-width division so the heavy discovery corpus gets the memory-budgeted spawn width (CI 24min->~10min), concurrent memory peak fail-closed under budget (#5421) * WIP: CI scheduler fix: model per-node resource demand + stop the executor's e * CI scheduler: per-node memory demand + stop run_walk even-width division The discovery corpus (572 witnesses, ~1041s) ran SERIALLY at width=1: the host run_walk chunked each batch by spawn_width and split the budget evenly across chunk siblings (per_runnable_width = width / chunk.len()). The corpus was node 0 of a 4-node chunk -> width=1, while the 3 cheap gates it shared the chunk with idled (a SingleClaim ignores width entirely). 1000s tentpole on one thread. Host (claim_executor.rs): run_walk now launches the whole readiness-layer batch concurrently and hands every node the full spawn_width. A DiscoveryBatch shards up to that width (the budgeted .dag authority); a SingleClaim is one resolve thread and ignores it. No more even division -> the heavy corpus gets the full memory-budgeted width (=4 on the fleet). Model (ci_floor_plan.dag): per-node concurrent memory demand (units of the corpus per-shard peak). The corpus fills the memory budget at its width, so a co-running second heavy resolve (source-root ingest, ~one shard's peak) would exceed budget (4+1 > 4 units = 70 > 62.5 GiB) and OOM (exit 137). Derived a ResourceDependsOn edge serializing source-root ingest behind the corpus into its own readiness layer. ci_fleet exposes gunbc_ci_floor_memory_budget_units. Oracle (§5): gunbc_ci_floor_plan_fits_memory_budget asserts every scheduled readiness layer's concurrent peak <= budget; discriminating control proves the serialization edge is load-bearing (drop it -> corpus layer overflows -> RED). Witnesses updated: 3 readiness layers, source-ingest serialized, plan oracle. Expected CI ~24min -> ~10min (corpus ~1041s/4 + serialized ingest ~244s). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Merge origin/main into session/sharp-bee-349 Adapt to #5419 (memory-aware spawn-width removed, floor width pinned to 4): drop the re-added placement/budget-units model; keep the run_walk host fix and re-express OOM safety as a flat topology policy — at most one heavy whole-tree resolve per readiness layer (source_root_ingest serialized behind the corpus). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Brian Searls <briansrls@gunb.ai> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
briansrls
added a commit
that referenced
this pull request
Jun 21, 2026
…pinned 4) Wire the measurement->plan loop #5431 opened: gunbc_ci_floor_spawn_width now returns min(shard_count, hardware_threads, floor(0.8*live_budget / measured_peak)) instead of hardware_thread_count(4). - gunbc.ci_floor_measurement (new): committed MEASURED per-shard peak, stamped from #5431's VmHWM emit (run 27893441091: 8400113664 bytes at width 4 => /4 = ~2.1GB/shard), NOT the ~14GB hand-grounded literal #5419 deleted. + dissolve-on drift-gate marker. - std.realization_width: memory_aware_spawn_width fold (pure, v2-inheriting) with a 0.8 safety fraction; fail-closed to a conservative committed width when budget unreadable. - claim_executor: reads the LIVE cgroup memory.max (meminfo fallback) and threads it into the .dag fold via run_in_context_with_args (DESIGN §3 measured=>peripheral: we commit our measurement, read the external-authored budget live). - witnesses: non-binding (4->7), tight-budget backoff, unreadable->fallback, fraction load-bearing, sub-unit->1. Proven by execution: derived spawn_width=7. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
briansrls
added a commit
that referenced
this pull request
Jun 21, 2026
…erified) Memory-aware width model was removed as unwired (#5419); current floor width is a pinned constant 4 (ci_floor_plan.dag:292), ~3% of 128c. #5444 re-adds the memory term (not yet on main). The lever lifts both the pin and the shard_count cap from the envelope. Fresh spawn-width slice to be reopen-scoped by quick-ant. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
briansrls
added a commit
that referenced
this pull request
Jun 21, 2026
…quick-ant) The §3 invariant conflated two phases. spawn_width is the discovery-corpus RUN-phase shard width (witnesses against the prebuilt binary — cores∧mem-bound, pids-light), so it does NOT multiply per_build_pids; width-up is pids-safe on its own. The pids crash is the BUILD phase: concurrent_builds × per_build_pids ≤ pids_cap, coupling host-packing × fan-out, not spawn_width. Two invariants, not one product. Also: #5375 memory-aware was superseded by #5419's pin (#5444 re-adds the term). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
briansrls
added a commit
that referenced
this pull request
Jun 21, 2026
…r one authority) (#5462) * docs/plans: compute-envelope-model — one authority for the CI fleet's resource dimensions Plan doc resolving the §1 ROADMAP "CI on compute fabric" pointer. Models the bimodal crash-or-idle pathology as one root (N hand-tuned resource dimensions, no single ResourceEnvelope authority) and the §3 fix: derive every knob — spawn_width, fan-out, TasksMax, jobserver, MemoryMax — from one measured envelope, with the public(shape)/ctrl(realization) split. Co-owned warm-lark-306 + quick-ant-298 (§1 lead); CC bright-stag-194 (ROADMAP + test profile). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * compute-envelope-model: correct current width = pinned 4 (quick-ant verified) Memory-aware width model was removed as unwired (#5419); current floor width is a pinned constant 4 (ci_floor_plan.dag:292), ~3% of 128c. #5444 re-adds the memory term (not yet on main). The lever lifts both the pin and the shard_count cap from the envelope. Fresh spawn-width slice to be reopen-scoped by quick-ant. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ROADMAP §1: point the CI-on-fabric chore-line at the compute-envelope plan doc Makes #5462 self-contained (doc + its own pointer, atomic, no orphan window). Different line from the §1 nightly-reframe edits — no collision. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * compute-envelope-model: split RUN-phase width from BUILD-phase pids (quick-ant) The §3 invariant conflated two phases. spawn_width is the discovery-corpus RUN-phase shard width (witnesses against the prebuilt binary — cores∧mem-bound, pids-light), so it does NOT multiply per_build_pids; width-up is pids-safe on its own. The pids crash is the BUILD phase: concurrent_builds × per_build_pids ≤ pids_cap, coupling host-packing × fan-out, not spawn_width. Two invariants, not one product. Also: #5375 memory-aware was superseded by #5419's pin (#5444 re-adds the term). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ROADMAP §1: fix chore-line spawn_width conflation (RUN/BUILD split) main's compact bullet (#5459) said 'spawn_width (memory- & pids-aware)' — the same conflation quick-ant corrected in compute-envelope-model.md. spawn_width is RUN-phase (cores∧mem, prebuilt binary, pids-light); pids binds the BUILD phase (fan-out × concurrent-builds). Keeps ROADMAP consistent with the doc in one PR. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ROADMAP §1: adopt quick-ant's precise spawn_width de-conflation wording Explicit RUN/BUILD phase labels — spawn_width = cores∧mem (RUN, prebuilt binary, pids-light); pids burst belongs on per-build fan-out (BUILD, concurrent_builds × per-build-pids ≤ TasksMax). No "pids" on the spawn_width bullet at all (the conflation's verification check). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ROADMAP §1: shorten CI-on-fabric chore-line to a summary + doc pointer Per operator: ROADMAP lines stay scannable progress summaries; density goes in plan docs (its own stated rule — no restating detail). The 1175-char inline chore-list/diagnosis/formulas now live in compute-envelope-model.md; the line is a one-sentence summary + pointer. Density preserved in the doc, not lost. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Brian Searls <briansrls@gunb.ai> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
briansrls
added a commit
that referenced
this pull request
Jun 21, 2026
… peak + live budget (not pinned 4) (#5444) * WIP: §1-C memory-aware width (consume the merged #5431 measurement keystone): * §1-C: derive floor spawn_width from measured peak + live budget (not pinned 4) Wire the measurement->plan loop #5431 opened: gunbc_ci_floor_spawn_width now returns min(shard_count, hardware_threads, floor(0.8*live_budget / measured_peak)) instead of hardware_thread_count(4). - gunbc.ci_floor_measurement (new): committed MEASURED per-shard peak, stamped from #5431's VmHWM emit (run 27893441091: 8400113664 bytes at width 4 => /4 = ~2.1GB/shard), NOT the ~14GB hand-grounded literal #5419 deleted. + dissolve-on drift-gate marker. - std.realization_width: memory_aware_spawn_width fold (pure, v2-inheriting) with a 0.8 safety fraction; fail-closed to a conservative committed width when budget unreadable. - claim_executor: reads the LIVE cgroup memory.max (meminfo fallback) and threads it into the .dag fold via run_in_context_with_args (DESIGN §3 measured=>peripheral: we commit our measurement, read the external-authored budget live). - witnesses: non-binding (4->7), tight-budget backoff, unreadable->fallback, fraction load-bearing, sub-unit->1. Proven by execution: derived spawn_width=7. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fmt: rustfmt claim_executor §1-C changes Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * WIP: §1-C memory-aware width (consume the merged #5431 measurement keystone): * §1-C review fix: type byte fields as std.measure ByteSize (REQUEST_CHANGES) Addresses claude-opus-4-7 REQUEST_CHANGES: the public byte-carrying fields in dsl/std/realization_width.dag (the authority layer) must consume std.measure ByteSize, not flat Int (DESIGN §2/§3 unit modeling). - memory_aware_spawn_width + memory_aware_width_value: memory_budget / per_shard_peak are now ByteSize (the canonical Memory carrier). Counts projected to Int at the boundary via byte_size_count_int (Nat subset Int, the int_min/cores widening idiom). - gunbc.ci_floor_measurement.per_shard_peak_rss_bytes returns ByteSize. - ci_floor_plan grounds the raw host Int budget into byte_size() at the host seam. - internal arithmetic helpers stay Int on extracted counts (not flagged; the placement_supply idiom). Fixes the prior WIP's Int/Nat if-branch type error. Proven by execution: realization_width_witnesses + ci_floor_plan_witnesses green; end-to-end claim_executor still derives spawn_width=7. compile --target rust clean (only the pre-existing unrelated gcp utf8_decode_bytes error remains). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * WIP: §1-C memory-aware width (consume the merged #5431 measurement keystone): * WIP: §1-C memory-aware width (consume the merged #5431 measurement keystone): * §1-C: de-anemic measured peak — typed ConcurrentMemoryPeakSample + measure-op per-share derivation Operator review: the floor's measured peak was modeled as two loose data rows (peak: ByteSize + spawn_width: Int) with the per-shard division inlined in the consumer (byte_size(count: byte_size_count(...) / int)) — anemic on two axes: - the observation is a FACT BUNDLE (a peak RSS only meaningful AT the concurrency it was sampled at: concurrent peak ≈ per-share × concurrency) held as two rows that can drift apart and neither reads as 'a measurement' (DESIGN §2/§3); - the per-share derivation was a one-off unwrap/divide/re-ground expression, not a grounded measure operation (DESIGN §4 — operations come from inhabitance). Fix, modeled on the existing std measure-algebra (time_measure_seq/_list_total in std.realization_measurement, which project via measure_count → magnitude arithmetic → reconstruct): - std.realization_measurement: add the agnostic SHAPE — type ConcurrentMemoryPeakSample { peak: ByteSize, concurrency: Int } + concurrent_memory_peak_per_share(sample) that divides the Memory magnitude by the sample concurrency and re-grounds at the same carrier; fail-closed (§5) on non-positive concurrency → whole peak (share count 1). + discriminating witness witness_concurrent_memory_peak_per_share_divides, wired into the keystone test aggregator. - gunbc.ci_floor_measurement: the measured VALUES now live as ONE typed sample (peak 8400113664 @ concurrency 4, §3 measured⇒peripheral); per-shard fn delegates to the std measure op — no inline arithmetic. Green by execution: keystone witnesses true; gunbc_ci_floor_per_shard_peak_rss_bytes() = ByteSize { count: 2100028416 } (= 8400113664 / 4, unchanged); realization_width_witnesses true. Consumer signature (ci_floor_plan) unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Brian Searls <briansrls@gunb.ai> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The memory-aware spawn-width / placement-width / memory-budget model turned out to be elaborate
spec that nothing live consumed except the CI floor's concurrency — which is a trivial constant
(#5390 literally pinned "width=4"). This rips it out and keeps CI green on a flat width.
Cut
gunbc_ci_floor_spawn_width()→hardware_thread_count(4)(claim_executor stillevaluates the companion fn — "host never decides width"). Dropped the memory-budget oracle.
gunbc_ci_runner_spec(theruns-onsource forci.yml).placement_spawn_width/placement_memory_width_bound/placement_floor_memory_budget_bytes. Kept the feasibility projection (placement_supply_row,offer_placement_supply_row),satisfies, and all the types (used by other witnesses).Kept / added
gunbc/operator_fleet.dag(new): srv1/srv2 as plainComputeOfferinventory — real hardware,BMC credentials as never-committed
std.credentialshandles, operator-owned. The authorityCI/sessions can derive from when a real consumer needs it; no placement modeling attached.
Verification (DESIGN §5)
0 diagnostics; the v2 floor closure compiles;ci_floor_plan_witnessesruns GREENon the constant width — the live CI floor is unaffected.
Supersedes #5390 (allocator) and #5417 — both can be closed.
🤖 Generated with Claude Code