Skip to content

CI floor ledger correction (§9): the cold-child class — corrected batch mechanisms, lever-1 retired, per-gate pooled child is the new rank-1 - #7123

Closed
briansrls wants to merge 4 commits into
mainfrom
claude/ci-floor-child-ledger
Closed

briansrls wants to merge 4 commits into
mainfrom
claude/ci-floor-child-ledger

Conversation

@briansrls

Copy link
Copy Markdown
Contributor

Appends the Pi/srv1 probe session's findings to the merged #7106 ledger as dated rows (snapshots stand; this is an amendment, not a retraction). Mechanism verified in-tree before writing: run_gunbc_claims (dag/tools/host_prelude.dag) folds over claims spawning one cold gunbc run --claim-run child per ClaimRun — serial, no short-circuit — each rebuilding the module index and closure resolve from scratch. ~26 children/run ≈ 23 of the 48 floor minutes, invisible to resolves_total by design.

  • Corrected mechanisms: batch-1's wall is the sum of 12 serial children (Σ=480s), not a parallel-gate max; batch-6's ingest bin costs 0.008s (the 12.1-min wall was 12 children — the "bin re-walks the tree" story is retracted as mechanism); batch-7 is 2 children.
  • Counterfactual receipts (srv1, PR-head binaries): single cold child 47.2s; 12 claims pooled in one process 93.4s vs CI's 480s; ingest pooled 106.5s vs 190.5s → per-gate pooling displaces 13–17 min/run. One tree, one resolve: floor stages consume the compile-clean computation (levers 2/3/8 of #7106) #7122's post-fix residuals corroborate: the gap between them and the counterfactuals is the remaining child tax.
  • Lever 1 retired as priced: Floor time: dissolve the #6848 once-per-entry fixpoint + reconcile-miss qualified-fill (~18min residual) #7030 + floor: precompute both-closure edges to dissolve per-entry bare-ref fixpoint #7056 are already ancestors of the measured run with zero recovery (971ms/group persists; the denominator ignored 1738/2206 affected-set skips). Re-diagnose before spending.
  • Fix shape with the memory coupling stated: pooled-per-gate child now (one-function blast radius, separate process that dies and frees); executor-warm index sharing only with the eviction lane — the executor is the 16 GiB-pinned crawl process of run 29976854620.
  • Pi exaggeration-bench receipts (cold 2-entry resolve >25 min vs ~2s warm) and the workflow-lane process lesson (receipt freshness applies at the diagnosis grain).

Remaining Pi probes (pooled variants, discovery n-scaling, whole-tree resolve) append here as they land.

🤖 Generated with Claude Code

https://claude.ai/code/session_016fdkaGGLUKpLRwwqxp5sLg


Generated by Claude Code

…t never named — corrected batch mechanisms, lever-1 retired as stale, per-gate pooled child is the new rank-1

Appends the Pi/srv1 probe session's corroborated-and-corrected findings to the
merged #7106 ledger (dated rows, no retraction of the snapshots):

- NEW CLASS: cold-index-per-process — run_gunbc_claims (dag/tools/
  host_prelude.dag) forks one cold gunbc child per ClaimRun, serial, no
  short-circuit; ~26 children/run = ~23 of 48 floor minutes (48%), invisible
  to resolves_total by design. Third and largest instance of the one root
  (#7030 per-thread, double-resolve-rewire per-entry, this per-process).
- Corrected mechanisms: batch-1 wall is the SUM of 12 serial children (not
  parallel max); batch-6's ingest bin costs 0.008s (the wall was 12 children;
  the 'bin re-walks tree' story retracted); batch-7 is 2 children.
- Counterfactual receipts (srv1, PR-head bins): single child 47.2s; 12 pooled
  in one process 93.4s vs CI's 480s; ingest pooled 106.5s vs 190.5s. Per-gate
  pooling displaces 13-17 min/run. #7122's residuals are the remaining child
  tax (4.65 vs ~1.6; 5.24 vs ~1.8).
- Lever 1 RETIRED as priced: #7030+#7056 are ancestors of the measured run
  with zero recovery; 971ms/group persists; denominator ignored 1738/2206
  affected-set skips. Re-diagnose before spending.
- Fix shape with the memory coupling stated: pooled-per-gate child now
  (separate process, dies and frees); executor-warm sharing only with the
  eviction lane (the executor is the 16GiB-pinned crawl process).
- Pi exaggeration receipts + the workflow-lane process lesson (receipt
  freshness at the diagnosis grain).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016fdkaGGLUKpLRwwqxp5sLg
@cursor

cursor Bot commented Jul 23, 2026

Copy link
Copy Markdown

Bugbot is not enabled for your account, so this pull request was not reviewed.

Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs.

claude added 3 commits July 23, 2026 12:55
…agent history sweep, receipts throughout)

docs/plans/resolve-regression-journey.md answers the operator's 2026-07-23
question ('what fundamentally keeps regressing? we have fixed resolve
several times and it comes back worse') from a 65-event dated fix ledger
over origin/main since 2026-06-25 plus the in-tree receipt docs, with an
adversarial verification pass:

- The measured trajectory: ~8m floor (pre-flip) -> discovery flip x12
  demand (#6438, 2026-07-09) -> timeout bounce 30..270 raised to fit ->
  #6848 (+18m time AND the 1GB/process parse baseline, RSS 6.5->20GB,
  2026-07-20) -> mechanism-correct follow-ups recovering less than priced
  (M1 ~0% on capped hosts) -> the 4h memory.high crawl -> #7120/#7122/#9.
  Three mechanisms wear one trend line: added demand, retention to the
  cap, cap lost.
- The five grains of one duplication (per-thread/-entry/-run/-process/-PR),
  discovered serially, fixed independently, no shared already-computed
  authority — instance-patching as validation where construction
  (ComputationIdentity) is the fix.
- Verified: NO floor-time regression gate exists (the 5s law is enforced
  but scoped to eval; the regression mass lives in the exempted infra
  carve-out); 5+ merges since 07-01 added corpus-denominated work with
  zero merge-time cost pricing.
- What ends it, in order: the cost wall (budget refusal + regression
  gate), the identity authority, the retention lane, cost pricing at
  merge.

Registered in doc_graph_roots (reachability suite 7/7 PASS by execution).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016fdkaGGLUKpLRwwqxp5sLg
…teardown-tail + selection-control levers, the 2652-error two-surface hygiene divergence, Pi/srv stretch heuristic

Appends the paired Pi-vs-srvN decomposition's findings and QUALIFIES this
branch's own section-9 fix-shape claim:

- Pooled-child win has a memory cliff (Pi: pooled 73m ~= spawn-sum 66m;
  2.7GB unioned closure thrashes where 12 small children do not) —
  corroborated by 7122's post-rework residuals under slot pressure
  (3.08/6.90 vs 1.6/1.8 idle counterfactuals). Pooling pays fully only
  with per-worker index shrink or K sub-pools sized to the slot.
- Teardown tail ~2.5-3.1min (twice-confirmed: Pi swap-growth during Drop
  + 7122 R4 post-b7 teardown) — new lever row, fast-exit shape.
- Selection-control step 4m51s had no ledger row — new lever row.
- Whole-tree baselines: strict resolve 23m48s/31.4GB (~7x compile-clean);
  compile-clean 86% reconcile.
- Standing red flag (section-5 class, not a lever): bare gunbc compile on
  the CI-green tree exits 2,652 unlisted-import errors — the floor receipt
  path tolerates hygiene the CLI enforces; same class as the
  fleet_converge_emit standalone failures and import-strip Class B; needs
  a single hygiene authority.
- Bench heuristic recorded: Pi/srv stretch ~8x = CPU-shaped, 25-78x =
  memory-shaped.

Journey doc section-5 item 2 qualified accordingly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016fdkaGGLUKpLRwwqxp5sLg
…25+1), plumbing-PR cost profile (only per-batch residuals are honest on plumbing PRs), green-run cgroup pegging receipt, per-call vs per-gate pooling gap

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016fdkaGGLUKpLRwwqxp5sLg
@briansrls briansrls closed this Jul 23, 2026
briansrls added a commit that referenced this pull request Jul 23, 2026
* CI floor endgame: per-batch claim pool, teardown fast-exit, selection-control shared index, discovery phase receipts, THE COST WALL, wet-lane re-homes (D1-D6)

D1: batch-0's claim-backed gates share ONE pooled claim_batch child
(tools.cheap_gate_pool — the union of the two concern authorities,
ByDerivation; K sub-pools sized to the slot via
cheap_gate_pool_max_claims_per_child, the section-9.1 cliff wall; child
stays a separate process). New CheapClaimPoolGate runs it; layering/
extdeps gates became non-vacuous enrollment walls asserting
chunk-transform totality (bare pool membership would be a tautology).
Ingest's 4 mktemp overlay children stay — the genuine constraint, named
in-row (ingest_pool_separation_note).

D2: floor_terminal_fast_exit after receipts flush skips the 2.5-3.1min
Drop walk of the ~16GB retained store (twice-confirmed). Exit code
preserved (walk_exit_code, unit-pinned); truncated receipts still red
(unwritable_receipt_base_reds_not_vanishes RED control). Terminal path
only.

D3: the selection-control step's 4m51s was three cold whole-pool index
builds inside floor_skip_discovery_witness; the warmup case now rides
resolve_entry_graph_shared (one shared build + the deliberately-cold
Class-B control build, named). Step stays per-PR; ledger row added.

D4: discovery pump phases land as typed rows in the floor resolve
receipt (discovery_pump_wall_ms / roster_walk / diff_observe /
frontier_attribution / shared_index_build / preresolve_calibration /
runner_resolve + corpus resolve/eval serial sums). Lever-1 stays
retired as priced; the fresh profile names the next mechanism before
any spend.

D5 THE COST WALL: per-batch wall budgets as data
(gunbc.ci_spec.gunbc_ci_floor_batch_wall_budget_seconds, sum 53min
under the 55min cap; per-batch never per-run — the plumbing-PR
profile), read fail-closed at arm time, enforced as typed
FLOOR-BATCH-OVER-BUDGET refusals (never a widen), recorded as typed
receipt rows (target/floor-batch-wall-receipt.txt). Ruling
reconciliation in the carrier note (budgets are admission/scheduling,
never witness verdicts — the 5s-eval-law split; endgame brief operator
is the sign-off). Raise requires an appended receipt note. RED control
both directions proven by execution on the tighten-only injection
fixture (budget_red_control_plan.dag).

D6: bin_witness_wet_per_row_wall_budget_seconds=60; the two rows over
it (floor_skip keystone 289s — also a per-PR duplicate of the
selection-control step — and cross_shard_seam live-tree 183s) re-home
to falsifier wet lane 5 as typed frontier rows
(falsifier_rehomed_bin_wet_rows, reasons + dissolve_on, registry
enrolled); 52 rows remain the per-PR smoke subset.

Also: carries #7123's docs baseline (journey + attribution 9-9.2, now
linked into the doc graph), appends attribution 9.3 with the <=30min
post-merge prediction, retires completed lever rows, appends the
cost-wall landing to the timeout-note bounce history, and heals
ci.yml's stale fleet-converge.sh heal-line (its artifact and
registration were deleted by #7121; the committed yml predated the
regen).

Proven by execution: executor battery 18/18; budget gate RED (witness
PASSes, batch refusal fires, exit 1, OverBudget receipt row) and GREEN
(exit 0, WithinBudget) on the fixture plan; ONE claim_batch child ran
all 4 batch-0 gates with the ONE nested 12-claim pooled child (3 PASS +
drift PASS post-regen); chunk-totality witnesses 4/4 incl. the RED
discriminator; ci_floor_plan_witnesses PASS (membership 4, falsifier
5 lanes, budget coverage); whole-tree compile: zero new error
identities vs main (main itself reds 2,647 pre-existing unlisted-import
rows — the flagged non-goal); doc-reachability wall green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Rc3TWb8FFexdVQDNbb44

* Merge main + heal ROADMAP.md trailing-newline drift on the merged tree

The PR's first floor run failed exactly one gate: generated_artifact_drift
on ROADMAP.md — main's committed copy (post-#7118) carries a trailing
blank line its own generator does not emit, so the merge tree drifted.
Regenerated via main_wet on the merged authorities; ci.yml converged
byte-identical with main's auto-heal of the fleet-converge line this
branch had healed independently. Budget-coverage witness re-proven true
on the merged tree (7 batches, 7 budget rows).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Rc3TWb8FFexdVQDNbb44

* D6 follow-through: FalsifierRehomedBinWet consumer cadence — the re-homed rows name their executing consumer (Phase 0(b))

The a9744d3 floor red was WITNESS ADMISSION REFUSAL
cause=UnexecutedDeferredWitness count=2: the D6 re-home removed the two
rows from bin_witness_wet_entries but the admission machinery had no
concept of the new falsifier lane, so the excluded witnesses read as
enrolled-with-zero-consumers — the invariant working as built.

Wired through every reader: std.witness_admission gains the
FalsifierRehomedBinWet cadence variant; ci_layer_roots re-exports it,
flips floor_skip_discovery_witness_test.dag's exclusion row onto it
(with its own reason/dissolve strings; the seam file stays
BinWitnessWet — its algebra row still backs it), and adds the
witness_exclusion_row_is_rehomed_bin_wet predicate (all 7 existing
cadence matches gain the arm); v2.workflow.witness_admission routes
consumer_for_explicit_rosters and the explicit-consumer manifest
through falsifier_rehomed_bin_wet_entries and adds rehomed backing to
witness_exclusion_explicit_roster_rows_consistent (new param);
cli_run's source-scan classifier gains the RehomedBinWetRow head and
the classification roster gains the name; the refusal message names
the lane.

Proven by execution: reconciliation witnesses 5/5 green incl. the NEW
RED control (an unbacked FalsifierRehomedBinWet row reds);
witness_admission_invariant_holds green with the two new rows (the
keystone row classifies FalsifierRehomedBinWet and the manifest covers
it); Rust unit test pins the RehomedBinWetRow head parse.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Rc3TWb8FFexdVQDNbb44

* Operator sign-off round: budget-note signature + follow-up rows, D2 comment scope fix, 9.3 process lesson + on-call notes

- gunbc_ci_floor_batch_wall_budget_note gains the dated operator
  signature (briansrls 2026-07-23) affirming the admission/verdict
  reading — the declared human-signed exemption row the 2026-07-10
  ruling requires — and the amended raise discipline (raising needs a
  dated operator-signed line naming run id + enrollment; tightening by
  ordinary receipt note; remedy is diagnose-or-signed-raise, never
  rerun), plus three named follow-up rows: the prelude coverage hole
  (~5min before batch-1 arms is outside every budget), identity-keyed
  budgets (index coupling rides the ComputationIdentity lane), and the
  K=16 union-RSS receipt (provisional until the pooled child's
  per-shard-peak-rss lands in a CI log; local proxy 1.82GB).
- floor_terminal_fast_exit doc corrected: it is the common tail of
  run()'s walk path (all plan walks), not floor-only.
- Attribution 9.3 gains the process lesson (every piece proven by
  execution except one full floor walk — the lane that redded, twice)
  and the on-call pre-positioning (batches 3/6 at 1.29x/1.45x headroom
  inside slot noise; first organic over-budget expected within days;
  diagnose-or-signed-raise, never pre-widen).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Rc3TWb8FFexdVQDNbb44

* Admission head-scan: word-boundary guard — a head substring inside a longer identifier is not a roster row

Run 30033250697 panicked (exit 101, fail-closed as designed) because the
raw substring scan matched bin_wet( inside the NEW predicate name
witness_exclusion_row_is_rehomed_bin_wet(row: ...) — no entry: literal
in the window, so the scanner stopped the line. The class fix: a head
occurrence preceded by an identifier char is a longer name, skipped
before parsing. Unit-pinned both ways
(head_scan_ignores_longer_identifiers_containing_a_head; the live-file
witness_admission_deferred_rows_have_consumers test now covers the exact
CI path — it sat in the locally OOM-killed battery, which is how this
escaped: the section-9.3 process lesson at unit grain).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Rc3TWb8FFexdVQDNbb44

* Census re-admission: regenerate witness_entry_eligibility_census for cheap_gate_pool_test (864 -> 865)

Run 30035504886's exit-101 was #7127's census refusal working correctly:
the floor's witness_execution_leg_label refuses any discovery entry with
no census TSV row, and this branch's new dag/test/claim/
cheap_gate_pool_test.dag postdated the committed census. Regenerated via
scripts/witness_entry_eligibility_census.sh (the emit transport itself
refused at 865 roster vs 864 declared — fail-closed both ways),
witness_entry_eligibility_census_entry_count bumped 864 -> 865, and the
three .dag sync witnesses proven green by execution through claim_batch
(declared-count exposed, carrier paths, tsv sync). The Rust sync test is
CWD-fragile by construction (relative witness_layer_roots under cargo's
package-dir cwd — the same pre-existing local-suite class as the two
scope-test failures reproduced on pristine main); the .dag witnesses are
the CI-executing consumers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Rc3TWb8FFexdVQDNbb44

* Heal census-count hand-pin drift: witness literal + note 876 -> 880

fbd55eb's floor red was the census_count witness (from main's #7127) —
witness_entry_eligibility_census_count_holds hardcodes the entry count
as a literal (== 876) SEPARATE from the data constant this branch
already bumped to 880, so 880 == 876 was false. The tsv-sync witness
could not catch it (it checks declared == TSV-rows, both statically 880).

Verified the true count is 880, not an emit artifact: tree clean (zero
untracked test files), independent git-tracked find of *_test.dag with
test fn/data = 880, TSV data rows = 880, cheap_gate_pool_test present.
Decomposition: main's committed 876 is STALE — main's live find is 879
(git ls-tree find on origin/main), undetected because affected-set
selection skips this witness on PRs that do not touch its import
closure; this plumbing PR's affected set is the first to run it against
the drifted tree. So 880 = main's true 879 + this PR's cheap_gate_pool
(+1), and updating the literal HEALS main's 3-file latent census drift
in passing.

The count lives in four hand-synced copies (this literal, its note,
witness_entry_eligibility_census_entry_count, the committed TSV) — a §3
parallel-representation main's carrier keeps in sync by hand; restructure
is out of this PR's scope, noted in the witness note. Proven by
execution: all census witnesses across census_test / classification_test
/ tsv_sync_test green through claim_batch (count_holds, the six
classification rows incl. bulk-pending default for the new entry, tsv
sync).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Rc3TWb8FFexdVQDNbb44

---------

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants