Skip to content

compute-fabric: cut the unwired spawn-width/placement over-modeling; pin floor width=4; keep fleet inventory - #5419

Merged
briansrls merged 1 commit into
mainfrom
compute-fabric-cleanup
Jun 21, 2026
Merged

briansrls merged 1 commit into
mainfrom
compute-fabric-cleanup

Conversation

@briansrls

Copy link
Copy Markdown
Contributor

The memory-aware spawn-width / placement-width / memory-budget model turned out to be elaborate
spec that nothing live consumed except the CI floor's concurrency — which is a trivial constant
(#5390 literally pinned "width=4"). This rips it out and keeps CI green on a flat width.

Cut

  • ci_floor_plan: gunbc_ci_floor_spawn_width() → hardware_thread_count(4) (claim_executor still
    evaluates the companion fn — "host never decides width"). Dropped the memory-budget oracle.
  • ci_fleet: dropped the spawn-width/placement fns + their imports; kept the offer + the live
    gunbc_ci_runner_spec (the runs-on source for ci.yml).
  • compute_fabric: deleted placement_spawn_width / placement_memory_width_bound /
    placement_floor_memory_budget_bytes. Kept the feasibility projection (placement_supply_row,
    offer_placement_supply_row), satisfies, and all the types (used by other witnesses).
  • ci_floor_plan_witness_test: dropped the 4 spawn-width witnesses; kept the 8 plan-structure ones.

Kept / added

  • gunbc/operator_fleet.dag (new): srv1/srv2 as plain ComputeOffer inventory — real hardware,
    BMC credentials as never-committed std.credentials handles, operator-owned. The authority
    CI/sessions can derive from when a real consumer needs it; no placement modeling attached.

Verification (DESIGN §5)

  • dsl compile 0 diagnostics; the v2 floor closure compiles; ci_floor_plan_witnesses runs GREEN
    on the constant width — the live CI floor is unaffected.

Supersedes #5390 (allocator) and #5417 — both can be closed.

🤖 Generated with Claude Code

…tion; pin floor width=4

The memory-aware spawn-width / placement-width / memory-budget model was elaborate spec that
nothing live consumed except the CI floor's concurrency — and that is a trivial constant. Cut it:

- ci_floor_plan: gunbc_ci_floor_spawn_width -> hardware_thread_count(4) (claim_executor still
  evaluates it; host never decides width). Drop the memory-budget oracle fn.
- ci_fleet: drop the spawn-width/placement fns + their imports; keep the offer + the live
  gunbc_ci_runner_spec (the runs-on source for ci.yml).
- compute_fabric: delete placement_spawn_width / placement_memory_width_bound /
  placement_floor_memory_budget_bytes (the width parallelization). Keep the feasibility
  projection (placement_supply_row / offer_placement_supply_row) + satisfies + the types.
- ci_floor_plan_witness_test: drop the 4 spawn-width witnesses; keep the plan-structure ones.

Adds gunbc/operator_fleet.dag — srv1/srv2 as plain ComputeOffer inventory (real hardware; BMC
creds as never-committed std.credentials handles; operator-owned). The authority CI/sessions
derive from when a real consumer needs it; no placement modeling attached.

Verified: dsl compile 0-diagnostics; the v2 floor closure compiles and ci_floor_plan_witnesses
runs GREEN on the constant width. Supersedes #5390 (allocator) + #5417.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@briansrls
briansrls merged commit 6164261 into main Jun 21, 2026
1 check passed
@briansrls
briansrls deleted the compute-fabric-cleanup branch June 21, 2026 01:20
briansrls added a commit that referenced this pull request Jun 21, 2026
Adapt to #5419 (memory-aware spawn-width removed, floor width pinned to 4):
drop the re-added placement/budget-units model; keep the run_walk host fix and
re-express OOM safety as a flat topology policy — at most one heavy whole-tree
resolve per readiness layer (source_root_ingest serialized behind the corpus).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
briansrls added a commit that referenced this pull request Jun 21, 2026
…s even-width division so the heavy discovery corpus gets the memory-budgeted spawn width (CI 24min->~10min), concurrent memory peak fail-closed under budget (#5421)

* WIP: CI scheduler fix: model per-node resource demand + stop the executor's e

* CI scheduler: per-node memory demand + stop run_walk even-width division

The discovery corpus (572 witnesses, ~1041s) ran SERIALLY at width=1: the
host run_walk chunked each batch by spawn_width and split the budget evenly
across chunk siblings (per_runnable_width = width / chunk.len()). The corpus
was node 0 of a 4-node chunk -> width=1, while the 3 cheap gates it shared the
chunk with idled (a SingleClaim ignores width entirely). 1000s tentpole on one
thread.

Host (claim_executor.rs): run_walk now launches the whole readiness-layer
batch concurrently and hands every node the full spawn_width. A DiscoveryBatch
shards up to that width (the budgeted .dag authority); a SingleClaim is one
resolve thread and ignores it. No more even division -> the heavy corpus gets
the full memory-budgeted width (=4 on the fleet).

Model (ci_floor_plan.dag): per-node concurrent memory demand (units of the
corpus per-shard peak). The corpus fills the memory budget at its width, so a
co-running second heavy resolve (source-root ingest, ~one shard's peak) would
exceed budget (4+1 > 4 units = 70 > 62.5 GiB) and OOM (exit 137). Derived a
ResourceDependsOn edge serializing source-root ingest behind the corpus into
its own readiness layer. ci_fleet exposes gunbc_ci_floor_memory_budget_units.

Oracle (§5): gunbc_ci_floor_plan_fits_memory_budget asserts every scheduled
readiness layer's concurrent peak <= budget; discriminating control proves the
serialization edge is load-bearing (drop it -> corpus layer overflows -> RED).
Witnesses updated: 3 readiness layers, source-ingest serialized, plan oracle.

Expected CI ~24min -> ~10min (corpus ~1041s/4 + serialized ingest ~244s).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Merge origin/main into session/sharp-bee-349

Adapt to #5419 (memory-aware spawn-width removed, floor width pinned to 4):
drop the re-added placement/budget-units model; keep the run_walk host fix and
re-express OOM safety as a flat topology policy — at most one heavy whole-tree
resolve per readiness layer (source_root_ingest serialized behind the corpus).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Brian Searls <briansrls@gunb.ai>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
briansrls added a commit that referenced this pull request Jun 21, 2026
…pinned 4)

Wire the measurement->plan loop #5431 opened: gunbc_ci_floor_spawn_width now
returns min(shard_count, hardware_threads, floor(0.8*live_budget / measured_peak))
instead of hardware_thread_count(4).

- gunbc.ci_floor_measurement (new): committed MEASURED per-shard peak, stamped from
  #5431's VmHWM emit (run 27893441091: 8400113664 bytes at width 4 => /4 = ~2.1GB/shard),
  NOT the ~14GB hand-grounded literal #5419 deleted. + dissolve-on drift-gate marker.
- std.realization_width: memory_aware_spawn_width fold (pure, v2-inheriting) with a 0.8
  safety fraction; fail-closed to a conservative committed width when budget unreadable.
- claim_executor: reads the LIVE cgroup memory.max (meminfo fallback) and threads it into
  the .dag fold via run_in_context_with_args (DESIGN §3 measured=>peripheral: we commit
  our measurement, read the external-authored budget live).
- witnesses: non-binding (4->7), tight-budget backoff, unreadable->fallback, fraction
  load-bearing, sub-unit->1. Proven by execution: derived spawn_width=7.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
briansrls added a commit that referenced this pull request Jun 21, 2026
…erified)

Memory-aware width model was removed as unwired (#5419); current floor width
is a pinned constant 4 (ci_floor_plan.dag:292), ~3% of 128c. #5444 re-adds the
memory term (not yet on main). The lever lifts both the pin and the shard_count
cap from the envelope. Fresh spawn-width slice to be reopen-scoped by quick-ant.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
briansrls added a commit that referenced this pull request Jun 21, 2026
…quick-ant)

The §3 invariant conflated two phases. spawn_width is the discovery-corpus
RUN-phase shard width (witnesses against the prebuilt binary — cores∧mem-bound,
pids-light), so it does NOT multiply per_build_pids; width-up is pids-safe on
its own. The pids crash is the BUILD phase: concurrent_builds × per_build_pids
≤ pids_cap, coupling host-packing × fan-out, not spawn_width. Two invariants,
not one product. Also: #5375 memory-aware was superseded by #5419's pin (#5444
re-adds the term).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
briansrls added a commit that referenced this pull request Jun 21, 2026
…r one authority) (#5462)

* docs/plans: compute-envelope-model — one authority for the CI fleet's resource dimensions

Plan doc resolving the §1 ROADMAP "CI on compute fabric" pointer. Models the
bimodal crash-or-idle pathology as one root (N hand-tuned resource dimensions,
no single ResourceEnvelope authority) and the §3 fix: derive every knob —
spawn_width, fan-out, TasksMax, jobserver, MemoryMax — from one measured
envelope, with the public(shape)/ctrl(realization) split. Co-owned warm-lark-306
+ quick-ant-298 (§1 lead); CC bright-stag-194 (ROADMAP + test profile).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* compute-envelope-model: correct current width = pinned 4 (quick-ant verified)

Memory-aware width model was removed as unwired (#5419); current floor width
is a pinned constant 4 (ci_floor_plan.dag:292), ~3% of 128c. #5444 re-adds the
memory term (not yet on main). The lever lifts both the pin and the shard_count
cap from the envelope. Fresh spawn-width slice to be reopen-scoped by quick-ant.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ROADMAP §1: point the CI-on-fabric chore-line at the compute-envelope plan doc

Makes #5462 self-contained (doc + its own pointer, atomic, no orphan window).
Different line from the §1 nightly-reframe edits — no collision.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* compute-envelope-model: split RUN-phase width from BUILD-phase pids (quick-ant)

The §3 invariant conflated two phases. spawn_width is the discovery-corpus
RUN-phase shard width (witnesses against the prebuilt binary — cores∧mem-bound,
pids-light), so it does NOT multiply per_build_pids; width-up is pids-safe on
its own. The pids crash is the BUILD phase: concurrent_builds × per_build_pids
≤ pids_cap, coupling host-packing × fan-out, not spawn_width. Two invariants,
not one product. Also: #5375 memory-aware was superseded by #5419's pin (#5444
re-adds the term).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ROADMAP §1: fix chore-line spawn_width conflation (RUN/BUILD split)

main's compact bullet (#5459) said 'spawn_width (memory- & pids-aware)' — the
same conflation quick-ant corrected in compute-envelope-model.md. spawn_width is
RUN-phase (cores∧mem, prebuilt binary, pids-light); pids binds the BUILD phase
(fan-out × concurrent-builds). Keeps ROADMAP consistent with the doc in one PR.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ROADMAP §1: adopt quick-ant's precise spawn_width de-conflation wording

Explicit RUN/BUILD phase labels — spawn_width = cores∧mem (RUN, prebuilt
binary, pids-light); pids burst belongs on per-build fan-out (BUILD,
concurrent_builds × per-build-pids ≤ TasksMax). No "pids" on the spawn_width
bullet at all (the conflation's verification check).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ROADMAP §1: shorten CI-on-fabric chore-line to a summary + doc pointer

Per operator: ROADMAP lines stay scannable progress summaries; density goes in
plan docs (its own stated rule — no restating detail). The 1175-char inline
chore-list/diagnosis/formulas now live in compute-envelope-model.md; the line
is a one-sentence summary + pointer. Density preserved in the doc, not lost.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Brian Searls <briansrls@gunb.ai>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
briansrls added a commit that referenced this pull request Jun 21, 2026
… peak + live budget (not pinned 4) (#5444)

* WIP: §1-C memory-aware width (consume the merged #5431 measurement keystone):

* §1-C: derive floor spawn_width from measured peak + live budget (not pinned 4)

Wire the measurement->plan loop #5431 opened: gunbc_ci_floor_spawn_width now
returns min(shard_count, hardware_threads, floor(0.8*live_budget / measured_peak))
instead of hardware_thread_count(4).

- gunbc.ci_floor_measurement (new): committed MEASURED per-shard peak, stamped from
  #5431's VmHWM emit (run 27893441091: 8400113664 bytes at width 4 => /4 = ~2.1GB/shard),
  NOT the ~14GB hand-grounded literal #5419 deleted. + dissolve-on drift-gate marker.
- std.realization_width: memory_aware_spawn_width fold (pure, v2-inheriting) with a 0.8
  safety fraction; fail-closed to a conservative committed width when budget unreadable.
- claim_executor: reads the LIVE cgroup memory.max (meminfo fallback) and threads it into
  the .dag fold via run_in_context_with_args (DESIGN §3 measured=>peripheral: we commit
  our measurement, read the external-authored budget live).
- witnesses: non-binding (4->7), tight-budget backoff, unreadable->fallback, fraction
  load-bearing, sub-unit->1. Proven by execution: derived spawn_width=7.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fmt: rustfmt claim_executor §1-C changes

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* WIP: §1-C memory-aware width (consume the merged #5431 measurement keystone):

* §1-C review fix: type byte fields as std.measure ByteSize (REQUEST_CHANGES)

Addresses claude-opus-4-7 REQUEST_CHANGES: the public byte-carrying fields in
dsl/std/realization_width.dag (the authority layer) must consume std.measure ByteSize,
not flat Int (DESIGN §2/§3 unit modeling).

- memory_aware_spawn_width + memory_aware_width_value: memory_budget / per_shard_peak
  are now ByteSize (the canonical Memory carrier). Counts projected to Int at the
  boundary via byte_size_count_int (Nat subset Int, the int_min/cores widening idiom).
- gunbc.ci_floor_measurement.per_shard_peak_rss_bytes returns ByteSize.
- ci_floor_plan grounds the raw host Int budget into byte_size() at the host seam.
- internal arithmetic helpers stay Int on extracted counts (not flagged; the
  placement_supply idiom). Fixes the prior WIP's Int/Nat if-branch type error.

Proven by execution: realization_width_witnesses + ci_floor_plan_witnesses green;
end-to-end claim_executor still derives spawn_width=7. compile --target rust clean
(only the pre-existing unrelated gcp utf8_decode_bytes error remains).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* WIP: §1-C memory-aware width (consume the merged #5431 measurement keystone):

* WIP: §1-C memory-aware width (consume the merged #5431 measurement keystone):

* §1-C: de-anemic measured peak — typed ConcurrentMemoryPeakSample + measure-op per-share derivation

Operator review: the floor's measured peak was modeled as two loose data rows
(peak: ByteSize + spawn_width: Int) with the per-shard division inlined in the
consumer (byte_size(count: byte_size_count(...) / int)) — anemic on two axes:

  - the observation is a FACT BUNDLE (a peak RSS only meaningful AT the concurrency
    it was sampled at: concurrent peak ≈ per-share × concurrency) held as two rows
    that can drift apart and neither reads as 'a measurement' (DESIGN §2/§3);
  - the per-share derivation was a one-off unwrap/divide/re-ground expression, not a
    grounded measure operation (DESIGN §4 — operations come from inhabitance).

Fix, modeled on the existing std measure-algebra (time_measure_seq/_list_total in
std.realization_measurement, which project via measure_count → magnitude arithmetic
→ reconstruct):

  - std.realization_measurement: add the agnostic SHAPE — type ConcurrentMemoryPeakSample
    { peak: ByteSize, concurrency: Int } + concurrent_memory_peak_per_share(sample) that
    divides the Memory magnitude by the sample concurrency and re-grounds at the same
    carrier; fail-closed (§5) on non-positive concurrency → whole peak (share count 1).
    + discriminating witness witness_concurrent_memory_peak_per_share_divides, wired into
    the keystone test aggregator.
  - gunbc.ci_floor_measurement: the measured VALUES now live as ONE typed sample
    (peak 8400113664 @ concurrency 4, §3 measured⇒peripheral); per-shard fn delegates to
    the std measure op — no inline arithmetic.

Green by execution: keystone witnesses true; gunbc_ci_floor_per_shard_peak_rss_bytes()
= ByteSize { count: 2100028416 } (= 8400113664 / 4, unchanged); realization_width_witnesses
true. Consumer signature (ci_floor_plan) unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Brian Searls <briansrls@gunb.ai>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant