From cf0db444eb820bc05abf87972d579bb3ceb61e62 Mon Sep 17 00:00:00 2001 From: Brian Searls Date: Thu, 23 Jul 2026 04:46:05 +0000 Subject: [PATCH 1/7] docs: attribute the CI floor 45-72min band by ci-job phase and batch. Log-diff receipts on fleet runs decompose workflow vs ci-job wall time and name discovery resolve, self-host gates, and #7030 effectful recovery as the dominant buckets. Co-authored-by: Cursor --- .../ci-floor-time-45-72-band-attribution.md | 181 ++++++++++++++++++ 1 file changed, 181 insertions(+) create mode 100644 docs/plans/ci-floor-time-45-72-band-attribution.md diff --git a/docs/plans/ci-floor-time-45-72-band-attribution.md b/docs/plans/ci-floor-time-45-72-band-attribution.md new file mode 100644 index 00000000000..a7fefe915ee --- /dev/null +++ b/docs/plans/ci-floor-time-45-72-band-attribution.md @@ -0,0 +1,181 @@ +# CI floor time audit — attribute the 45–72 min band + +**Status:** measurement receipt, 2026-07-23 (session vivid-fox-471). **DESIGN.md + carriers remain +authority** — this doc is a timestamped profiling receipt, not a fact ledger. Dissolves when +`realization_measurement_loop` Phase-0 lands a durable `.dag`-native Gantt carrier that supersedes +prose receipts (same trigger as [ci-floor-fractal-gantt.md](ci-floor-fractal-gantt.md)). + +**One-line verdict:** The **ci job** (not whole workflow) clusters **44–55 min** on current main; +the **45–72 min workflow band** is mostly **build + ci + deploy_dashboard**. Inside the ci job, +**~35 min** is `claim_executor` floor time, dominated by **discovery resolve (~12 min wall, +~620s serial-sum)** and **self-host gate chain batches 5–6 (~14 min combined)**. The #6848 +namespace-walk regression (+24 min vs PRE) is **partially recovered** on effectful gates (#7030, +−11 min) but **discovery resolve (+520s vs PRE)** and **new self-host enrollment (+14 min)** +keep the band ~25–30 min above the Jul-20 PRE baseline. + +--- + +## 1. Band census (what “45–72 min” measures) + +Measured over the last 60 `ci.yml` runs on **main** (2026-07-22/23 window). + +| Grain | n | min | p25 | median | p75 | max | 45–72 band | +|---|---:|---:|---:|---:|---:|---:|---| +| **Workflow wall** (build+ci+deploy) | 52 | 23m | 37m | 55m | 64m | 98m | 29/52 (56%) | +| **ci job only** | 52 | 3m* | 20m | 44m | 48m | 55m | 21/52 (40%) | + +\*Floor of 3m = fast-lane / early-fail runs. + +**Interpretation:** Operator-facing “CI takes ~an hour” is the **workflow** number (median 55m). +Performance work should track the **ci job** (median 44m, p75 48m) — build is ~1m and deploy is +~2m and should not pollute floor attribution. + +**Not the same as ~72 min emit:** `gunbc_ci_witness_corpus_only_batches_note` cites **~72 min** for +the whole-tree `--target dag` compile-clean gate as **pre-push infeasible** scope — that is +**emit wall inside `dag_compile_clean_gate_passes` on a cold whole-tree closure**, not the ci job +wall. On current main the compile-clean **batch wall is ~1 min** because CI runs import-closure +scoped compile (`tools.dag_compile_clean_scope`), not whole-tree emit every PR. + +--- + +## 2. ci job decomposition (receipt runs) + +Log-diff harness: parse `claim_executor` invocations by +`--plan-function gunbc_ci_regen_floor_batches` vs `gunbc_ci_floor_batches`; batch walls from +`claim_executor: batch N` → next batch start (or last `PASS [batch N]`). + +| arm | run | ci job | overhead† | regen | floor | cap-sat | +|---|---|---:|---:|---:|---:|---| +| PRE-#6848 | `29763408563` | **17.9m** | ~7m | 2.5m | 10.8m | no | +| POST-#6848 | `29819122813` | **42.1m** | ~11m | 3.6m | 28.9m | no | +| POST-#6998/#6999 | `29855080611` | — | — | — | 28.9m‡ | **yes** | +| POST-#7030 | `29880571548` | — | — | 3.4m | 36.9m | no | +| main Jul-23 | `29970583893` | **47.7m** | ~10m | 3.0m | 35.0m | no | +| main Jul-23 | `29967907137` | **49.2m** | ~12m | 3.3m | 36.0m | no | +| main Jul-22 wide | `29961193892` | **48.7m** | ~10m | 3.0m | 36.5m | no | + +†Overhead = ci job wall minus regen + floor executor spans (unpack release bins, plan prelude +~99s, artifact verify, yaml/selection prelude, post-floor steps). +‡Failed run; batch-4 effectful wall still ~19m (pre-#7030 class). + +**Current main typical ci job (~48m) ≈ 10m overhead + 3m regen + 35m floor.** + +--- + +## 3. Floor executor batch attribution (current schedule) + +Schedule after #7030 (heavy-resolve main-thread routing) + self-host gate enrollment. Representative: +run `29970583893` (ci job 47.7m, success). + +| batch | lane | wall | % of floor | dominant mechanism | +|---:|---|---:|---:|---| +| 0 | cheap gates (layering / extdeps / drift) | <1m† | <3% | negligible scans (#7088 moves these before compile-clean on newer main) | +| 1 | `dag_compile_clean_gate_passes` | **0.8m** | 2% | import-closure scoped `.dag` compile (not whole-tree emit) | +| 2 | discovery corpus (hermetic, SelectionApplied) | **12.2m** | **35%** | **entry resolve** — see §4.1 | +| 3 | wet corpora (exec + bin witnesses) | **1.3m** | 4% | small explicit roster; resolve ~11s serial | +| 4 | effectful gates (emit_host + cheap scans) | **7.1m** | 20% | host effects; **recovered** from ~18.5m by #7030 | +| 5 | `source_root_ingest` + self-host chain | **10.9m** | **31%** | heavy whole-tree resolve + ingest host work | +| 6 | `self_host_reads_real_bytes` | **2.9m** | 8% | heavy resolve + filesystem read gate | +| — | regen sub-plan (separate invocation) | **3.0m** | — | regen_verify ~2m + staleness ~1m | + +†On `29970583893` cheap gates still co-reside in batch 4; post-#7088 they move to batch 0 +(seconds). + +### 3.1 Discovery `[measurement]` line (batch 2) + +| arm | resolve serial | eval serial | witnesses | skipped | +|---|---:|---:|---:|---:| +| PRE `29763408563` | **99s** | 15s | 2087 | 1652 | +| POST `29819122813` | **403s** | 15s | 2128 | 1755 | +| main `29970583893` | **620s** | 15s | 2204 | 1739 | + +**Eval is not the story** on affected-set PRs (SelectionApplied skips ~79% of roster). **Resolve +serial-sum grew 6.3×** PRE→main and tracks batch-2 wall (~12 min). Per-witness amortized resolve +rose ~48ms → ~282ms (2204 witnesses), matching the #6848 bare-reference fixpoint class priced in +[floor-time-namespace-walk-regression-diagnosis.md](floor-time-namespace-walk-regression-diagnosis.md). + +### 3.2 Governor / cap saturation + +On **16 GiB `memory.high` capped hosts**, post-#6848 floors can run `forced_serial=1` with +`hard_backoffs=1` while RSS sits at 93–100% of budget (§1.4 of namespace-walk diagnosis). +This is an **additive throttle** on top of walk work — same resolve class, slower wall. Uncapped +hosts (MemAvailable budget) complete the same schedule without backoffs but **do not** erase the +resolve-serial inflation. + +--- + +## 4. Mechanism ledger (cause → minutes → owner) + +| # | mechanism | Δ vs PRE (~18m ci) | Δ vs POST-#6848 (~42m ci) | status | owner lane | +|---|---|---:|---:|---|---| +| A | #6848 bare-reference fixpoint (`extend_sources_to_both_closure_fixpoint`) — once per entry | **+5 min** floor batch-2 | ~0 (still dominant) | **open** | namespace-resolution §PR-5b | +| B | #6848 qualified-fill on reconcile miss | **+13 min** floor batch-4 (pre-#7030) | **−11 min** (#7030 heavy-resolve routing) | **partial** | #7030 landed; reconcile-miss path still open | +| C | `thread_local` cold `MultiEntryIndex` per spawned gate (#6999 didn't fix) | (folded into B) | **−11 min** | **landed #7030** | proud-bear-438 | +| D | Self-host gate enrollment batches 5–6 (`source_root_ingest`, `reads_real_bytes`) | **+14 min** | +14 min (new scope) | **by design** | self-host / module-identity | +| E | Cap-saturation throttle on 16 GiB runners | unpriced additive | widens band toward timeouts | **open** | v1-run-stability / governor envelope | +| F | `UnlistedImportUse` advisory generation (~5k rows/run) | typecheck overhead | same | **open** | namespace advisory suppression | +| G | ci overhead (bin unpack + plan prelude ~99s) | +3m | stable ~10m | baseline | CI substrate | + +**Reconciliation to band:** PRE ci 17.9m → main ci 47.7m ≈ **+30 min** ≈ A (+5) + B net (+2 after +#7030) + D (+14) + G (+3) + E (variable). The POST-#6848 → main delta is mostly **D** (new gates) +with **B partially reversed**. + +--- + +## 5. Historical contrast (effectful-gate recovery receipt) + +Batch-4 effectful wall (first gate `emit_host_gate_passes` → last batch-4 pass): + +| arm | batch-4 wall | notes | +|---|---:|---| +| PRE `29763408563` | **5.7m** | parallel groups, no self-host reads-bytes | +| POST `29819122813` | **18.5m** | cold per-thread index (#6848 + spawn routing) | +| POST `29855080611` | **18.8m** | #6998/#6999: ~0% recovery (by execution) | +| POST `29880571548` | **10.0m** | #7030: heavy-resolve flip — partial | +| main `29970583893` | **7.1m** | #7030 + schedule churn; near PRE+host-effects | + +--- + +## 6. Reproduction + +```bash +# ci job duration +gh run view --json jobs | jq '.jobs[] | select(.name=="ci") | {startedAt,completedAt,conclusion}' + +# floor batch walls + discovery measurement +gh run view --log > /tmp/floor.log +rg 'claim_executor: batch|PASS \[batch|\[measurement\] discovery corpus:|\[governor\] receipt:' /tmp/floor.log + +# compare arms +for r in 29763408563 29819122813 29970583893; do + echo "=== $r ===" + gh run view $r --json jobs | jq -r '.jobs[]|select(.name=="ci")|"ci job: \(.startedAt) -> \(.completedAt)"' + rg 'discovery corpus: [0-9]+ witness' <(gh run view $r --log 2>/dev/null) | head -1 +done +``` + +Python segmenter used for §2–3 tables: partition log on +`claim_executor" --source-root` + `--plan-function gunbc_ci_regen_floor_batches|gunbc_ci_floor_batches`. + +--- + +## 7. Next scoped dispatch (not this lane) + +Ordered by displaced ci job minutes on current main: + +1. **Discovery resolve fixpoint (A)** — attack once-per-entry bare-reference walk + (namespace-resolution-design §PR-5b residual). Target: **−5 to −8 min** floor batch-2. +2. **Self-host gate cost (D)** — per-gate measurement with `GUNBC_FLOOR_GANTT=1`; justify or + narrow batches 5–6 if host work is redundant with batch-4 emit_host. Target: clarity first; + reduction only if duplicated resolve. +3. **Cap saturation (E)** — headroom vs walk RSS on 16 GiB slots; pairs with #6848 residual. +4. **Advisory row suppression (F)** — if still hot after A. + +--- + +## 8. Provenance + +- Log-diff + ci job JSON: vivid-fox-471, by execution on fleet runs listed in §2–3 (2026-07-23). +- Parent mechanism docs: [floor-time-namespace-walk-regression-diagnosis.md](floor-time-namespace-walk-regression-diagnosis.md), + [ci-floor-fractal-gantt.md](ci-floor-fractal-gantt.md), PR #7030 receipt. +- Related open threads: DESIGN.md floor shared-computation memoization M2; namespace-resolution §PR-5b. From bc6dec3ad15206c34afca1fb082cff4e7080276a Mon Sep 17 00:00:00 2001 From: Brian Searls Date: Thu, 23 Jul 2026 04:49:13 +0000 Subject: [PATCH 2/7] docs: CI floor redundant-work ledger, phase TSV, and ranked levers. Re-derive run 29976989996 and five comparison arms; name per-stage duplicate work (per-entry resolve walks, ingest re-scan, cheap gates after compile) and price top levers in minutes without floor behavior changes. Co-authored-by: Cursor --- .../ci-floor-time-45-72-band-attribution.md | 243 ++++++++---------- .../ci_floor_lever_ranking_2026-07-23.tsv | 9 + .../ci_floor_phase_attribution_2026-07-23.tsv | 7 + ..._redundancy_ledger_skeleton_2026-07-23.tsv | 12 + 4 files changed, 140 insertions(+), 131 deletions(-) create mode 100644 docs/probes/ci_floor_lever_ranking_2026-07-23.tsv create mode 100644 docs/probes/ci_floor_phase_attribution_2026-07-23.tsv create mode 100644 docs/probes/ci_floor_redundancy_ledger_skeleton_2026-07-23.tsv diff --git a/docs/plans/ci-floor-time-45-72-band-attribution.md b/docs/plans/ci-floor-time-45-72-band-attribution.md index a7fefe915ee..245140d6470 100644 --- a/docs/plans/ci-floor-time-45-72-band-attribution.md +++ b/docs/plans/ci-floor-time-45-72-band-attribution.md @@ -1,181 +1,162 @@ -# CI floor time audit — attribute the 45–72 min band +# CI floor time audit — redundant-work ledger + lever ranking **Status:** measurement receipt, 2026-07-23 (session vivid-fox-471). **DESIGN.md + carriers remain -authority** — this doc is a timestamped profiling receipt, not a fact ledger. Dissolves when -`realization_measurement_loop` Phase-0 lands a durable `.dag`-native Gantt carrier that supersedes -prose receipts (same trigger as [ci-floor-fractal-gantt.md](ci-floor-fractal-gantt.md)). - -**One-line verdict:** The **ci job** (not whole workflow) clusters **44–55 min** on current main; -the **45–72 min workflow band** is mostly **build + ci + deploy_dashboard**. Inside the ci job, -**~35 min** is `claim_executor` floor time, dominated by **discovery resolve (~12 min wall, -~620s serial-sum)** and **self-host gate chain batches 5–6 (~14 min combined)**. The #6848 -namespace-walk regression (+24 min vs PRE) is **partially recovered** on effectful gates (#7030, -−11 min) but **discovery resolve (+520s vs PRE)** and **new self-host enrollment (+14 min)** -keep the band ~25–30 min above the Jul-20 PRE baseline. +authority** — prose + TSV receipts only; **no floor behavior changes** in this PR. Dissolves when +`realization_measurement_loop` Phase-0 lands a durable `.dag`-native Gantt carrier. ---- +**Product (operator mandate):** phase attribution is the **map**; the **product** is a per-stage +**redundant-work ledger** (what each stage recomputes that an earlier stage already computed on +the same input content) plus a **ranked lever table** priced in displaced minutes. -## 1. Band census (what “45–72 min” measures) +**Carriers (this PR):** -Measured over the last 60 `ci.yml` runs on **main** (2026-07-22/23 window). +- [`docs/probes/ci_floor_phase_attribution_2026-07-23.tsv`](../probes/ci_floor_phase_attribution_2026-07-23.tsv) — per-run per-phase walls +- [`docs/probes/ci_floor_redundancy_ledger_skeleton_2026-07-23.tsv`](../probes/ci_floor_redundancy_ledger_skeleton_2026-07-23.tsv) — stage × recomputes × duplicate-of +- [`docs/probes/ci_floor_lever_ranking_2026-07-23.tsv`](../probes/ci_floor_lever_ranking_2026-07-23.tsv) — ranked levers -| Grain | n | min | p25 | median | p75 | max | 45–72 band | -|---|---:|---:|---:|---:|---:|---:|---| -| **Workflow wall** (build+ci+deploy) | 52 | 23m | 37m | 55m | 64m | 98m | 29/52 (56%) | -| **ci job only** | 52 | 3m* | 20m | 44m | 48m | 55m | 21/52 (40%) | +--- -\*Floor of 3m = fast-lane / early-fail runs. +## 1. Band census (map only) -**Interpretation:** Operator-facing “CI takes ~an hour” is the **workflow** number (median 55m). -Performance work should track the **ci job** (median 44m, p75 48m) — build is ~1m and deploy is -~2m and should not pollute floor attribution. +| Grain | median | 45–72 min band | notes | +|---|---:|---|---| +| Workflow (build+ci+deploy) | 55m | 56% of main runs | operator-facing "~1 hour" | +| **ci job** | **44m** | 40% | **use this for floor attribution** | +| Floor step (`gunbc ci` claim_executor) | ~35–48m | — | regen excluded (~3m) | -**Not the same as ~72 min emit:** `gunbc_ci_witness_corpus_only_batches_note` cites **~72 min** for -the whole-tree `--target dag` compile-clean gate as **pre-push infeasible** scope — that is -**emit wall inside `dag_compile_clean_gate_passes` on a cold whole-tree closure**, not the ci job -wall. On current main the compile-clean **batch wall is ~1 min** because CI runs import-closure -scoped compile (`tools.dag_compile_clean_scope`), not whole-tree emit every PR. +The **~72 min** figure in `gunbc_ci_witness_corpus_only_batches_note` is whole-tree **emit** +infeasible for pre-push — not typical ci job wall. Scoped PRs skip compile-clean emit entirely +(`compile-clean scope: skipped`). --- -## 2. ci job decomposition (receipt runs) +## 2. Receipt anchor — run `29976989996` (re-derived) -Log-diff harness: parse `claim_executor` invocations by -`--plan-function gunbc_ci_regen_floor_batches` vs `gunbc_ci_floor_batches`; batch walls from -`claim_executor: batch N` → next batch start (or last `PASS [batch N]`). +Branch `session/gentle-raven-495`, green, srv1-01, ci job **54.9 min**, floor step **~48.4 min**. +7-batch schedule (post-#7088 cheap-gate early batch). **Whole-tree compile-clean** because diff +had no shard intersection. -| arm | run | ci job | overhead† | regen | floor | cap-sat | -|---|---|---:|---:|---:|---:|---| -| PRE-#6848 | `29763408563` | **17.9m** | ~7m | 2.5m | 10.8m | no | -| POST-#6848 | `29819122813` | **42.1m** | ~11m | 3.6m | 28.9m | no | -| POST-#6998/#6999 | `29855080611` | — | — | — | 28.9m‡ | **yes** | -| POST-#7030 | `29880571548` | — | — | 3.4m | 36.9m | no | -| main Jul-23 | `29970583893` | **47.7m** | ~10m | 3.0m | 35.0m | no | -| main Jul-23 | `29967907137` | **49.2m** | ~12m | 3.3m | 36.0m | no | -| main Jul-22 wide | `29961193892` | **48.7m** | ~10m | 3.0m | 36.5m | no | +| phase | wall (min) | % of floor | +|---|---:|---:| +| preamble (plan resolve + hygiene) | 1.9 | 4% | +| compile-clean receipt (whole-tree emit) | 3.6 | 7% | +| batch 1 cheap gates (3 nodes, 1 resolve-group) | **10.1** | **21%** | +| batch 2 compile gate consume | 0.5 | 1% | +| batch 3 discovery (663 entry-groups, 2206 rows) | **12.8** | **26%** | +| batch 4 wet corpora | 1.3 | 3% | +| batch 5 emit_host | 0.1 | 0% | +| batch 6 source_root_ingest (ONE node) | **12.1** | **25%** | +| batch 7 reads_real_bytes | 3.3 | 7% | -†Overhead = ci job wall minus regen + floor executor spans (unpack release bins, plan prelude -~99s, artifact verify, yaml/selection prelude, post-floor steps). -‡Failed run; batch-4 effectful wall still ~19m (pre-#7030 class). +**Top-3 = 35.0 of 48.4 min (72%):** discovery 12.8 + source_root_ingest 12.1 + cheap gates 10.1. -**Current main typical ci job (~48m) ≈ 10m overhead + 3m regen + 35m floor.** +Governor receipt: `budget=16GiB` (cgroup memory.high), `max_width_reached=1`, +`measured worker share=3.36GB`, `peak_current=10.1GiB`, `cross_worker_store withheld`. +Declared cold resolves: **4** (matches `ci_floor_declared_resolve_count`). --- -## 3. Floor executor batch attribution (current schedule) +## 3. Redundancy ledger (product) -Schedule after #7030 (heavy-resolve main-thread routing) + self-host gate enrollment. Representative: -run `29970583893` (ci job 47.7m, success). +Each row: what the stage computes, what earlier stage already computed on the **same content**, +and redundancy class per DESIGN §2 (duplicated / unnecessary / irrelevant). -| batch | lane | wall | % of floor | dominant mechanism | -|---:|---|---:|---:|---| -| 0 | cheap gates (layering / extdeps / drift) | <1m† | <3% | negligible scans (#7088 moves these before compile-clean on newer main) | -| 1 | `dag_compile_clean_gate_passes` | **0.8m** | 2% | import-closure scoped `.dag` compile (not whole-tree emit) | -| 2 | discovery corpus (hermetic, SelectionApplied) | **12.2m** | **35%** | **entry resolve** — see §4.1 | -| 3 | wet corpora (exec + bin witnesses) | **1.3m** | 4% | small explicit roster; resolve ~11s serial | -| 4 | effectful gates (emit_host + cheap scans) | **7.1m** | 20% | host effects; **recovered** from ~18.5m by #7030 | -| 5 | `source_root_ingest` + self-host chain | **10.9m** | **31%** | heavy whole-tree resolve + ingest host work | -| 6 | `self_host_reads_real_bytes` | **2.9m** | 8% | heavy resolve + filesystem read gate | -| — | regen sub-plan (separate invocation) | **3.0m** | — | regen_verify ~2m + staleness ~1m | +| stage | recomputes | duplicate of | class | receipt | +|---|---|---|---|---| +| **compile-clean receipt** | whole-tree load + resolve + typecheck + emit | — (first whole-tree touch) | **necessary** | 3.6min; builds `process_shared_index` | +| **cheap gates (batch 1)** | re-resolve witness entry + scan imports/extdeps/drift | compile-clean receipt on **same** `witness_layer_roots` | **duplicated** | 10.1min **after** 3.6min compile; 3 gates parallel, same resolve-group | +| **compile gate consume** | reads receipt artifact | compile-clean receipt | **necessary** | 27s verify only | +| **discovery** | per-entry `extend_sources_to_both_closure_fixpoint` + eval | compile-clean typed cache **in principle**; **not** per-entry walk | **duplicated per-entry** | resolve serial **643s**; `reusing process_shared_index` but #6848 walk dominates | +| **source_root_ingest** | `discover_source_root_ingest` bin full tree scan | compile-clean + discovery on same roots | **duplicated** | **12.1min** one node; separate binary path | +| **reads_real_bytes** | heavy whole-tree resolve + filesystem read | prior heavy gates | **duplicated heavy resolve** | 3.3min; serial after ingest | +| **width=1 governor** | serializes all witness work | — | **irrelevant** (scheduling) | NOT proposing cap raise; index shrink / M2 lane | +| **materialization unkeyed** | 2.19M unkeyed pure calls | keyed memo path | **duplicated (identity unknown)** | unkeyed=47% of demand; ComputationIdentity lane | -†On `29970583893` cheap gates still co-reside in batch 4; post-#7088 they move to batch 0 -(seconds). +**Key finding vs "4–5× whole-tree re-ingest" hypothesis:** declared **cold resolve count = 4** +per run — NOT four independent whole-tree cold graphs. The band is **not** four full re-ingests; +it is **one** whole-tree compile + **many per-entry walks** inside the shared index (discovery +643s serial resolve on 663 groups ≈ **970ms/group**), plus **two 12-min single-node gates** that +re-touch the tree through different code paths (ingest bin, cheap-gate scans). -### 3.1 Discovery `[measurement]` line (batch 2) +Full skeleton: [`ci_floor_redundancy_ledger_skeleton_2026-07-23.tsv`](../probes/ci_floor_redundancy_ledger_skeleton_2026-07-23.tsv). -| arm | resolve serial | eval serial | witnesses | skipped | -|---|---:|---:|---:|---:| -| PRE `29763408563` | **99s** | 15s | 2087 | 1652 | -| POST `29819122813` | **403s** | 15s | 2128 | 1755 | -| main `29970583893` | **620s** | 15s | 2204 | 1739 | +### 3.1 Batch-1 internal (cheap gates) -**Eval is not the story** on affected-set PRs (SelectionApplied skips ~79% of roster). **Resolve -serial-sum grew 6.3×** PRE→main and tracks batch-2 wall (~12 min). Per-witness amortized resolve -rose ~48ms → ~282ms (2204 witnesses), matching the #6848 bare-reference fixpoint class priced in -[floor-time-namespace-walk-regression-diagnosis.md](floor-time-namespace-walk-regression-diagnosis.md). +All three gates (`layering_imports`, `extdeps_external_authority`, `generated_artifact_drift`) +PASS at the **same timestamp** — one resolve-group, wall = **max** of parallel gate evals, not sum. +Dominant cost is the **shared resolve + host-effect scan** of the gate witness closure (~10min), +not one gate beating the others in serial. Per-gate split requires `GUNBC_FLOOR_GANTT=1` on a +replay (follow-up, not this audit PR). -### 3.2 Governor / cap saturation +### 3.2 Batch-6 / source_root_ingest — why 12 min for one node? -On **16 GiB `memory.high` capped hosts**, post-#6848 floors can run `forced_serial=1` with -`hard_backoffs=1` while RSS sits at 93–100% of budget (§1.4 of namespace-walk diagnosis). -This is an **additive throttle** on top of walk work — same resolve class, slower wall. Uncapped -hosts (MemAvailable budget) complete the same schedule without backoffs but **do not** erase the -resolve-serial inflation. +Evidence from run `29976989996` log: batch 6 invokes `discover_source_root_ingest` repeatedly +(shell `test -x` preamble then long-running ingest). This is a **separate release binary**, not +the compile-clean receipt path. It re-derives source-root ingest facts from the live tree — +work **not** consumed from the typed store the compile-clean receipt populated. Same pattern on +scoped main runs: batch 5 **10.9min** (`29970583893`) even when compile-clean is **skipped**. --- -## 4. Mechanism ledger (cause → minutes → owner) +## 4. Quadratic hunt (partial — historical arm) -| # | mechanism | Δ vs PRE (~18m ci) | Δ vs POST-#6848 (~42m ci) | status | owner lane | -|---|---|---:|---:|---|---| -| A | #6848 bare-reference fixpoint (`extend_sources_to_both_closure_fixpoint`) — once per entry | **+5 min** floor batch-2 | ~0 (still dominant) | **open** | namespace-resolution §PR-5b | -| B | #6848 qualified-fill on reconcile miss | **+13 min** floor batch-4 (pre-#7030) | **−11 min** (#7030 heavy-resolve routing) | **partial** | #7030 landed; reconcile-miss path still open | -| C | `thread_local` cold `MultiEntryIndex` per spawned gate (#6999 didn't fix) | (folded into B) | **−11 min** | **landed #7030** | proud-bear-438 | -| D | Self-host gate enrollment batches 5–6 (`source_root_ingest`, `reads_real_bytes`) | **+14 min** | +14 min (new scope) | **by design** | self-host / module-identity | -| E | Cap-saturation throttle on 16 GiB runners | unpriced additive | widens band toward timeouts | **open** | v1-run-stability / governor envelope | -| F | `UnlistedImportUse` advisory generation (~5k rows/run) | typecheck overhead | same | **open** | namespace advisory suppression | -| G | ci overhead (bin unpack + plan prelude ~99s) | +3m | stable ~10m | baseline | CI substrate | +Fit: discovery `resolve_serial_s` vs `entry_groups` (logged per run). -**Reconciliation to band:** PRE ci 17.9m → main ci 47.7m ≈ **+30 min** ≈ A (+5) + B net (+2 after -#7030) + D (+14) + G (+3) + E (variable). The POST-#6848 → main delta is mostly **D** (new gates) -with **B partially reversed**. +| run | class | entry_groups | resolve_serial_s | ms/group | +|---|---|---:|---:|---:| +| `29763408563` | PRE-6848 | ~500† | 99 | ~200 | +| `29819122813` | POST-6848 | ~500† | 403 | ~800 | +| `29976989996` | deep-diff | **663** | **644** | **971** | +| `29970583893` | trivial-diff | **663** | **620** | **935** | ---- +†PRE runs lack `adaptive pool over N entry-groups` log line; groups estimated from witness count. -## 5. Historical contrast (effectful-gate recovery receipt) +**Reading:** ms/group grew **~5×** PRE→POST (#6848 bare-reference fixpoint) while group count +grew ~30% (2087→2206 witnesses). The premium is **superlinear in per-group walk cost**, not +merely corpus size growth. **Local ptrace** on the two ~12min single-node gates is **not yet +run** (this audit PR is measurement-only); candidates: `rc_map_insert`, typecheck-env inductive +duplication, s1_closure re-walk (named in mandate). -Batch-4 effectful wall (first gate `emit_host_gate_passes` → last batch-4 pass): +--- + +## 5. Mandate questions — answers -| arm | batch-4 wall | notes | -|---|---:|---| -| PRE `29763408563` | **5.7m** | parallel groups, no self-host reads-bytes | -| POST `29819122813` | **18.5m** | cold per-thread index (#6848 + spawn routing) | -| POST `29855080611` | **18.8m** | #6998/#6999: ~0% recovery (by execution) | -| POST `29880571548` | **10.0m** | #7030: heavy-resolve flip — partial | -| main `29970583893` | **7.1m** | #7030 + schedule churn; near PRE+host-effects | +| # | question | answer | +|---|---|---| +| 1 | What dominates each duration class? | **Trivial-diff (~48m ci):** discovery (~12m) + source_root_ingest (~11m) + effectful (~7m). **Deep-diff (+6m):** adds whole-tree compile-clean (+3.6m) + cheap gates (+10m when pre-compile ordering). No 127–159m green runs in last 500 workflow samples — operator class may be falsifier/cold-control or older fleet. | +| 2 | Why 12min for source_root_ingest? | Separate `discover_source_root_ingest` binary re-scans tree; does not consume compile-clean receipt. Batch-1 gates: parallel group, ~10min shared resolve — per-gate split needs GANTT replay. | +| 3 | How many whole-tree index rebuilds? | **1** explicit whole-tree compile emit + **4** declared cold resolves — but **663 per-entry walks** inside discovery on shared index. `fe_begin` RSS climbs 9.5→15.2 GiB across discovery despite index reuse. | +| 4 | Width=1 fleet-wide on 16GiB? | **Yes on measured runs:** `max_width_reached=1`, `cross_worker_store withheld`. Worker share ~3.4GB leaves headroom on paper but governor does not grow width (width_growths=0). Recovery = per-worker index shrink / M2, **not** cap raise. | +| 5 | #6848 / #6999 claims? | **Verified:** resolve_serial 99→644s (+545s) PRE→seed; #6999 **~0%** batch-wall recovery on comparable hosts (29855080611 vs 29819122813). Discovery loads each entry once per worker at width=1 — memo hits near zero on that path. | --- -## 6. Reproduction +## 6. Ranked levers -```bash -# ci job duration -gh run view --json jobs | jq '.jobs[] | select(.name=="ci") | {startedAt,completedAt,conclusion}' - -# floor batch walls + discovery measurement -gh run view --log > /tmp/floor.log -rg 'claim_executor: batch|PASS \[batch|\[measurement\] discovery corpus:|\[governor\] receipt:' /tmp/floor.log - -# compare arms -for r in 29763408563 29819122813 29970583893; do - echo "=== $r ===" - gh run view $r --json jobs | jq -r '.jobs[]|select(.name=="ci")|"ci job: \(.startedAt) -> \(.completedAt)"' - rg 'discovery corpus: [0-9]+ witness' <(gh run view $r --log 2>/dev/null) | head -1 -done -``` +See [`ci_floor_lever_ranking_2026-07-23.tsv`](../probes/ci_floor_lever_ranking_2026-07-23.tsv). Top +three by displaced minutes: -Python segmenter used for §2–3 tables: partition log on -`claim_executor" --source-root` + `--plan-function gunbc_ci_regen_floor_batches|gunbc_ci_floor_batches`. +1. **Per-entry bare-reference fixpoint** — 8–12 min (namespace §PR-5b) +2. **source_root_ingest re-walk** — 10–12 min (module-identity lane) +3. **Cheap-gate scan after whole-tree compile** — 5–10 min (#7088 ordering may shift; sleek-crane owns) ---- +Config-grade follow-ups (named, not landed here): `GUNBC_FLOOR_GANTT=1` on fleet for per-gate +split; ptrace on ingest + discovery for quadratic stacks. -## 7. Next scoped dispatch (not this lane) +--- -Ordered by displaced ci job minutes on current main: +## 7. Reproduction -1. **Discovery resolve fixpoint (A)** — attack once-per-entry bare-reference walk - (namespace-resolution-design §PR-5b residual). Target: **−5 to −8 min** floor batch-2. -2. **Self-host gate cost (D)** — per-gate measurement with `GUNBC_FLOOR_GANTT=1`; justify or - narrow batches 5–6 if host work is redundant with batch-4 emit_host. Target: clarity first; - reduction only if duplicated resolve. -3. **Cap saturation (E)** — headroom vs walk RSS on 16 GiB slots; pairs with #6848 residual. -4. **Advisory row suppression (F)** — if still hot after A. +```bash +gh run view RUN_ID --log | rg 'claim_executor: batch|PASS \[batch|compile-clean scope|adaptive pool|discovery corpus:|\[governor\] receipt|floor materialization|floor resolve count' +``` --- ## 8. Provenance -- Log-diff + ci job JSON: vivid-fox-471, by execution on fleet runs listed in §2–3 (2026-07-23). -- Parent mechanism docs: [floor-time-namespace-walk-regression-diagnosis.md](floor-time-namespace-walk-regression-diagnosis.md), - [ci-floor-fractal-gantt.md](ci-floor-fractal-gantt.md), PR #7030 receipt. -- Related open threads: DESIGN.md floor shared-computation memoization M2; namespace-resolution §PR-5b. +- vivid-fox-471, 2026-07-23, log-diff by execution on runs in TSV. +- Parent mandate: sharp-bee-290 msg_eae17a34 (redundancy ledger + quadratic hunt). +- Related: [floor-time-namespace-walk-regression-diagnosis.md](floor-time-namespace-walk-regression-diagnosis.md), + [floor-shared-compute-memoization.md](floor-shared-compute-memoization.md), + [v1-run-stability-throughline.md](v1-run-stability-throughline.md). diff --git a/docs/probes/ci_floor_lever_ranking_2026-07-23.tsv b/docs/probes/ci_floor_lever_ranking_2026-07-23.tsv new file mode 100644 index 00000000000..ef0625df280 --- /dev/null +++ b/docs/probes/ci_floor_lever_ranking_2026-07-23.tsv @@ -0,0 +1,9 @@ +rank lever mechanism expected_min_recovered risk owning_thread evidence +1 Per-entry bare-reference fixpoint once-per-entry (#6848 residual) each discovery entry-group pays ~140ms namespace walk even with warm process_shared_index; 663 groups × ~140ms ≈ 5–8min discovery alone 8–12 low (perf-only; semantics frozen) namespace-resolution §PR-5b floor-time-namespace-walk-regression; resolve_serial 99s→644s PRE→seed +2 source_root_ingest_gate tree re-walk single-node batch 6/5 costs 10–12min via discover_source_root_ingest bin independent of compile-clean receipt 10–12 medium (must preserve ingest semantics) module-identity vs storage 29976989996 batch-6 12.1min; 29970583893 batch-5 10.9min +3 cheap_gates batch before compile consume (pre-7088 schedule) or gate scan duplication 10min parallel resolve-group scans layering/extdeps/drift after whole-tree compile already ran 5–10 low if derived from compile receipt ci fail-fast #7088 (sleek-crane owns ordering) 29976989996 batch-1 10.1min AFTER 3.6min compile +4 Unkeyed materialization / missing ComputationIdentity 2.19M unkeyed calls per run; 47% of demand unkeyed — duplicate pure work invisible to memo unknown until keyed medium duplicate-work graph lens 29976989996 unkeyed=2193388 duplicated=71948 +5 Width=1 latch from 16GiB slot + 3.4GB worker share governor never admits width>1; cross_worker_store withheld; discovery serial 0–4 (only if index shrinks) high if raising caps 5886 projection / M2 memo seed governor receipt; NOT proposing cap raise per mandate +6 rc_map_insert quadratic (per-entry map growth) superlinear in module count during pool_qualified_fill / reconcile budget TBD until profiled low bold-crane-271 quadratic hunt — local ptrace on 12min gates pending +7 UnlistedImportUse advisory generation (~5k rows/run) typecheck constructs advisories during every resolve 1–3 TBD low namespace advisory suppression floor-time-namespace-walk §1.5 +8 Whole-tree compile-clean on deep-diff (scope widen) deep-diff triggers 3.6min emit not present on scoped PRs 3–4 on scoped PRs only tools.dag_compile_clean_scope 29976989996 whole-tree vs 29970583893 skipped diff --git a/docs/probes/ci_floor_phase_attribution_2026-07-23.tsv b/docs/probes/ci_floor_phase_attribution_2026-07-23.tsv new file mode 100644 index 00000000000..67edcd896ad --- /dev/null +++ b/docs/probes/ci_floor_phase_attribution_2026-07-23.tsv @@ -0,0 +1,7 @@ +run_id branch class ci_job_min preamble_min compile_clean_min compile_scope wall_cheap_gates_min wall_compile_gate_min wall_discovery_min wall_wet_min wall_emit_host_min wall_source_root_ingest_min wall_reads_real_bytes_min witnesses entry_groups resolve_serial_s eval_serial_s resolves_total max_width peak_current_gib worker_share_gib unkeyed_calls duplicated_keys wasted_ms fe_begin_count_disc +29763408563 main PRE-6848 17.9 0.8 whole-tree or scoped 0.8 3.4 0.9 5.7 2087 99.3 14.9 3 1 unreadable 4 +29819122813 main POST-6848 42.1 0.7 whole-tree 0.7 8.4 1.2 18.5 2128 402.7 14.9 1 unreadable 4 +29976989996 session/gentle-raven-495 deep-diff green 54.9 1.89 3.58 whole-tree baseline (no shard intersection) 10.15 0.45 12.80 1.30 0.12 12.15 3.30 2206 663 643.5 18.4 4 1 10.1 3.36 2193388 71948 23407 4 +29970583893 main trivial-diff green 47.7 1.89 0 skipped (no touched paths) 0.8 12.17 1.30 7.10 10.90 2.90 2204 663 619.7 15.1 3 0 unreadable 2413946 71948 23407 4 +29967907137 main trivial-diff green 49.2 0 skipped 0.8 12.25 1.30 7.60 11.10 2.90 2204 663 623.8 15.2 3 0 unreadable +29880571548 main POST-7030 51.0 0.9 whole-tree partial 0.9 10.10 1.40 10.00 11.20 3.20 2137 478.0 15.6 4 1 unreadable diff --git a/docs/probes/ci_floor_redundancy_ledger_skeleton_2026-07-23.tsv b/docs/probes/ci_floor_redundancy_ledger_skeleton_2026-07-23.tsv new file mode 100644 index 00000000000..507fd73b549 --- /dev/null +++ b/docs/probes/ci_floor_redundancy_ledger_skeleton_2026-07-23.tsv @@ -0,0 +1,12 @@ +stage what_this_stage_computes prior_stage_with_same_input content_identity_shared redundancy_class duplicate_of evidence_run evidence_note +prelude_plan_resolve plan entry resolve + naming-hygiene walk + output policy install none (first touch) n/a necessary — 29976989996 t+113s to governor armed; ~1.9min +compile_clean_receipt whole-tree --target dag: load all modules, resolve graph, typecheck, emit dag artifacts none (first whole-tree touch) process_shared_index roots keyed to --source-root duplicated downstream compile_clean builds typed_module cache consumed by later resolves IN PRINCIPLE; later stages still pay per-entry walk work 29976989996 3.6min emit; receipt ok=true; scope=whole-tree baseline +cheap_gates_batch1 layering_imports + extdeps_authority + generated_artifact_drift scans over corpus compile_clean_receipt SAME witness_layer_roots closure; same process_shared_index duplicated compile_clean already typechecked whole tree; gates re-resolve witness entry + scan imports/extdeps/drift 29976989996 10.1min wall (3 gates parallel in 1 resolve-group); all PASS same timestamp +compile_gate_consume dag_compile_clean_gate_passes reads receipt only compile_clean_receipt SAME receipt artifact necessary (verify only) — 29976989996 27s — consumes receipt, does not re-emit +discovery_corpus per-entry resolve: extend_sources_to_both_closure_fixpoint + pool_qualified_fill + witness eval for ~663 entry-groups compile_clean_receipt + cheap_gates SAME pool roots; reuses process_shared_index at width=1 duplicated per-entry walk whole-tree typed cache does NOT elide per-entry bare-reference fixpoint (#6848); resolve serial 643s despite index reuse 29976989996 12.8min wall; 663 groups/2206 rows; fe_begin x4 during batch; cross_worker_store withheld width=1 +wet_corpus_small 20+53 explicit witness rows (exec/bin) discovery_corpus SAME index partially duplicated second/third discovery batch on subset; resolve serial 12s+11s 29976989996 1.3min combined +emit_host emit-host MVP smoke (cargo build subset) discovery_corpus SAME index necessary (host effect) distinct host-effect work 29976989996 7s +source_root_ingest_gate discover_source_root_ingest binary: full source-root ingest scan compile_clean + discovery SAME dag+src/v2 roots duplicated 12.1min for ONE node; runs separate bin re-walking tree; not served from compile-clean receipt alone 29976989996 12.15min wall; shell invokes discover_source_root_ingest repeatedly +reads_real_bytes_gate self_host_realized_comparison filesystem read gate compile_clean + discovery + source_root_ingest heavy_whole_tree_resolve on main thread duplicated heavy resolve 3.3min; third heavy-resolve gate serial chain 29976989996 batch 7 after ingest; #7030 routes to shared index but still pays resolve walk +governor_width width=1 entire run due to 16GiB memory.high budget + 3.36GB measured worker share all stages n/a irrelevant (scheduling) forces serial discovery; cross_worker_store withheld — cannot amortize across workers 29976989996 max_width_reached=1; peak_current=10.1GiB; NOT cap-saturated (no forced_serial) +materialization_unkeyed 2.19M unkeyed pure calls in one run all keyed stages n/a duplicated (computation identity unknown) materialization receipt: unkeyed_calls=2193388 vs keyed=2413946 29976989996 duplicate-work / ComputationIdentity lane; 47.6% of keyed+unkeyed unaccounted From 43e8790e6b8e451396a5953e8fa61f609a959a76 Mon Sep 17 00:00:00 2001 From: Brian Searls Date: Thu, 23 Jul 2026 04:53:42 +0000 Subject: [PATCH 3/7] docs: register ci-floor-time audit plan in doc_graph_roots. Unblocks doc_reachability pre-push gate for the measurement-only audit PR. Co-authored-by: Cursor --- dag/gunbc/doc_graph_roots.dag | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/dag/gunbc/doc_graph_roots.dag b/dag/gunbc/doc_graph_roots.dag index 97840eb24d5..8a7a9920d68 100644 --- a/dag/gunbc/doc_graph_roots.dag +++ b/dag/gunbc/doc_graph_roots.dag @@ -107,6 +107,12 @@ data hand_authored_doc_binds: List = [ work: DeclarationRef { module_path: "v2.workflow.ci_floor_plan", decl_name: "gunbc_ci_floor_batches", field: WholeDeclaration }, dissolution: HasTrigger { text: "dissolves into a registered gunbc.plan.Plan row when its P1 lands" }, }, + HandAuthoredDocBind { + home: PlanDoc, + slug: "ci-floor-time-45-72-band-attribution", + work: DeclarationRef { module_path: "gunbc.ci_materialization", decl_name: "ci_floor_declared_resolve_count", field: WholeDeclaration }, + dissolution: HasTrigger { text: "dissolves when realization_measurement_loop Phase-0 Gantt carrier supersedes prose receipts (same trigger as ci-floor-fractal-gantt)" }, + }, ] fn hand_authored_doc_graph_roots() -> List { From 652ec6544fae5d886b17de29fa9a320635a635ad Mon Sep 17 00:00:00 2001 From: Brian Searls Date: Thu, 23 Jul 2026 06:02:36 +0000 Subject: [PATCH 4/7] WIP: CI floor time audit: attribute the 45-72min band --- src/v1/stage0/src/cli_run.rs | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/src/v1/stage0/src/cli_run.rs b/src/v1/stage0/src/cli_run.rs index 6f3e1cad6e3..3f8bb189701 100644 --- a/src/v1/stage0/src/cli_run.rs +++ b/src/v1/stage0/src/cli_run.rs @@ -22116,6 +22116,14 @@ const NON_FOLD_RESIDUE_ROSTER: &[&str] = &[ "src/v2/compiler/05_emit_orchestration.dag::orch_emit_let_step", "src/v2/lens/live_read_classification.dag::live_read_carrier_eq", "src/v2/lens/live_read_classification.dag::path_pattern_eq", + // 2026-07-23 backfill: #7088 (cheap-gate early batch) landed two Gate-membership + // predicates unrostered on main — one-special-variant dispatch (fold over the declared + // cheap-gate list, nested `match g { => true _ => found }` off-variant arms). + // Surfaced when this PR's doc_graph_roots.dag edit re-ran the whole-corpus nfr receipt + // (same masking class as the dated blocks above). Burns down when Gate membership is + // derived from the declared roster fold instead of hand-nested matches. + "src/v2/workflow/ci_floor_plan.dag::gate_in_cheap_floor_membership", + "src/v2/workflow/ci_floor_plan.dag::spec_enrolls_gate", ]; fn nfr_strip_comments(content: &str) -> String { From 4ff4d219bf04ae9184dcc0b925d1dd2f25a35136 Mon Sep 17 00:00:00 2001 From: Brian Searls Date: Thu, 23 Jul 2026 06:56:53 +0000 Subject: [PATCH 5/7] WIP: CI floor time audit: attribute the 45-72min band --- src/v1/stage0/src/cli_run.rs | 8 -------- 1 file changed, 8 deletions(-) diff --git a/src/v1/stage0/src/cli_run.rs b/src/v1/stage0/src/cli_run.rs index fe379408e07..8f84dbd17dc 100644 --- a/src/v1/stage0/src/cli_run.rs +++ b/src/v1/stage0/src/cli_run.rs @@ -22126,14 +22126,6 @@ const NON_FOLD_RESIDUE_ROSTER: &[&str] = &[ "src/v2/compiler/05_emit_orchestration.dag::orch_emit_let_step", "src/v2/lens/live_read_classification.dag::live_read_carrier_eq", "src/v2/lens/live_read_classification.dag::path_pattern_eq", - // 2026-07-23 backfill: #7088 (cheap-gate early batch) landed two Gate-membership - // predicates unrostered on main — one-special-variant dispatch (fold over the declared - // cheap-gate list, nested `match g { => true _ => found }` off-variant arms). - // Surfaced when this PR's doc_graph_roots.dag edit re-ran the whole-corpus nfr receipt - // (same masking class as the dated blocks above). Burns down when Gate membership is - // derived from the declared roster fold instead of hand-nested matches. - "src/v2/workflow/ci_floor_plan.dag::gate_in_cheap_floor_membership", - "src/v2/workflow/ci_floor_plan.dag::spec_enrolls_gate", ]; fn nfr_strip_comments(content: &str) -> String { From 861b355b4076f562d03cbb2a48f5656942601349 Mon Sep 17 00:00:00 2001 From: Brian Searls Date: Thu, 23 Jul 2026 07:01:04 +0000 Subject: [PATCH 6/7] =?UTF-8?q?fix(ci):=20drop=20duplicate=20NFR=20roster?= =?UTF-8?q?=20+=20merge=20#7114=20=E2=80=94=20restore=20docs-only=20PR=20d?= =?UTF-8?q?iff?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The layering batch-1 red on 652ec65 was caused by our redundant cli_run.rs NFR backfill (already landed on main as #7114), which forced a whole-tree compile-clean on .rs and a heavier batch-1 path. doc_graph_roots bind row checked: import-bearing file, module_path in string literal — no new layering edge (not 7080-class). PR diff vs main is now docs + doc_graph only. Co-authored-by: Cursor From 0d9c7a132b8846ba14151b77ffaea828ac2b9a8d Mon Sep 17 00:00:00 2001 From: Brian Searls Date: Thu, 23 Jul 2026 11:44:00 +0000 Subject: [PATCH 7/7] docs(audit): add PR7110 270m timeout extreme-tail receipt MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Run 29986954853 (one-file v1_deletion_plan.dag diff) killed at ci-step timeout-minutes:270 — corpus/infra-bound tail, not diff-size-driven. TSV row + plan §1.1 + lever 5 sharpen + new lever 9 (silent ceiling ride). Co-authored-by: Cursor --- .../ci-floor-time-45-72-band-attribution.md | 18 ++++++++++++++++++ .../ci_floor_lever_ranking_2026-07-23.tsv | 3 ++- .../ci_floor_phase_attribution_2026-07-23.tsv | 1 + 3 files changed, 21 insertions(+), 1 deletion(-) diff --git a/docs/plans/ci-floor-time-45-72-band-attribution.md b/docs/plans/ci-floor-time-45-72-band-attribution.md index 245140d6470..1a972174f8d 100644 --- a/docs/plans/ci-floor-time-45-72-band-attribution.md +++ b/docs/plans/ci-floor-time-45-72-band-attribution.md @@ -23,11 +23,29 @@ the same input content) plus a **ranked lever table** priced in displaced minute | Workflow (build+ci+deploy) | 55m | 56% of main runs | operator-facing "~1 hour" | | **ci job** | **44m** | 40% | **use this for floor attribution** | | Floor step (`gunbc ci` claim_executor) | ~35–48m | — | regen excluded (~3m) | +| **Extreme tail (action ceiling)** | **270m+** | 1 receipt | PR #7110 run `29986954853` — see §1.1 | The **~72 min** figure in `gunbc_ci_witness_corpus_only_batches_note` is whole-tree **emit** infeasible for pre-push — not typical ci job wall. Scoped PRs skip compile-clean emit entirely (`compile-clean scope: skipped`). +### 1.1 Extreme tail — run `29986954853` (PR #7110, plan-only) + +**Receipt:** PR #7110 @ `ad44c819c`, run `29986954853`, **killed at the ci-step +`timeout-minutes: 270` ceiling** (workflow wall ~281 min including setup). Diff is **one file** +(`dag/gunbc/v1_deletion_plan.dag` only) — the most trivial affected-set path; compile-clean +should be skipped. + +**Reading:** NOT diff-size-driven. Strong evidence the tail is **corpus-denominated + +infra-bound** (serial `width=1` + memory thrash on a bad-luck runner). When width latches at 1 +there is **no recovery arm** — the run rides silently to the action ceiling (DESIGN §5 +absorbing-fallback shape: corpus-denominated cost breaks the budget later, not a typed refusal +mid-flight). Phase breakdown unavailable: log rotated on re-queue at cancel; floor never emitted +final batch receipts. + +**Lever sharpen:** extends ranked lever 5 (width=1 latch) — add **fail-fast at action ceiling** +as a separate scheduling finding (270 min silent ride vs early typed refusal). + --- ## 2. Receipt anchor — run `29976989996` (re-derived) diff --git a/docs/probes/ci_floor_lever_ranking_2026-07-23.tsv b/docs/probes/ci_floor_lever_ranking_2026-07-23.tsv index ef0625df280..5398acdc15d 100644 --- a/docs/probes/ci_floor_lever_ranking_2026-07-23.tsv +++ b/docs/probes/ci_floor_lever_ranking_2026-07-23.tsv @@ -3,7 +3,8 @@ rank lever mechanism expected_min_recovered risk owning_thread evidence 2 source_root_ingest_gate tree re-walk single-node batch 6/5 costs 10–12min via discover_source_root_ingest bin independent of compile-clean receipt 10–12 medium (must preserve ingest semantics) module-identity vs storage 29976989996 batch-6 12.1min; 29970583893 batch-5 10.9min 3 cheap_gates batch before compile consume (pre-7088 schedule) or gate scan duplication 10min parallel resolve-group scans layering/extdeps/drift after whole-tree compile already ran 5–10 low if derived from compile receipt ci fail-fast #7088 (sleek-crane owns ordering) 29976989996 batch-1 10.1min AFTER 3.6min compile 4 Unkeyed materialization / missing ComputationIdentity 2.19M unkeyed calls per run; 47% of demand unkeyed — duplicate pure work invisible to memo unknown until keyed medium duplicate-work graph lens 29976989996 unkeyed=2193388 duplicated=71948 -5 Width=1 latch from 16GiB slot + 3.4GB worker share governor never admits width>1; cross_worker_store withheld; discovery serial 0–4 (only if index shrinks) high if raising caps 5886 projection / M2 memo seed governor receipt; NOT proposing cap raise per mandate +5 Width=1 latch from 16GiB slot + 3.4GB worker share governor never admits width>1; cross_worker_store withheld; discovery serial; PR7110 run 29986954853 rode 270m action ceiling on ONE-FILE plan-only diff — no recovery arm 0–4 (only if index shrinks); tail unbounded without fail-fast medium 5886 projection / M2 memo seed governor receipt; 29986954853 extreme tail; NOT proposing cap raise per mandate +9 Action-ceiling silent ride (270m timeout) ci.yml timeout-minutes:270 kills without typed floor refusal; worst case on record for trivial diff unknown (scheduling policy) low ci workflow / floor disposition 29986954853 PR7110 plan-only; DESIGN §5 corpus-denominated-later shape 6 rc_map_insert quadratic (per-entry map growth) superlinear in module count during pool_qualified_fill / reconcile budget TBD until profiled low bold-crane-271 quadratic hunt — local ptrace on 12min gates pending 7 UnlistedImportUse advisory generation (~5k rows/run) typecheck constructs advisories during every resolve 1–3 TBD low namespace advisory suppression floor-time-namespace-walk §1.5 8 Whole-tree compile-clean on deep-diff (scope widen) deep-diff triggers 3.6min emit not present on scoped PRs 3–4 on scoped PRs only tools.dag_compile_clean_scope 29976989996 whole-tree vs 29970583893 skipped diff --git a/docs/probes/ci_floor_phase_attribution_2026-07-23.tsv b/docs/probes/ci_floor_phase_attribution_2026-07-23.tsv index 67edcd896ad..b7e7b3316e5 100644 --- a/docs/probes/ci_floor_phase_attribution_2026-07-23.tsv +++ b/docs/probes/ci_floor_phase_attribution_2026-07-23.tsv @@ -5,3 +5,4 @@ run_id branch class ci_job_min preamble_min compile_clean_min compile_scope wall 29970583893 main trivial-diff green 47.7 1.89 0 skipped (no touched paths) 0.8 12.17 1.30 7.10 10.90 2.90 2204 663 619.7 15.1 3 0 unreadable 2413946 71948 23407 4 29967907137 main trivial-diff green 49.2 0 skipped 0.8 12.25 1.30 7.60 11.10 2.90 2204 663 623.8 15.2 3 0 unreadable 29880571548 main POST-7030 51.0 0.9 whole-tree partial 0.9 10.10 1.40 10.00 11.20 3.20 2137 478.0 15.6 4 1 unreadable +29986954853 session/sharp-bee-290-roadmap-converge PR7110 plan-only ONE-FILE timeout@270m-action-ceiling 270+ killed unknown killed-before-receipt plan-only scoped (v1_deletion_plan.dag only) unknown unknown unknown unknown unknown unknown unknown unknown unknown unknown unknown 1? unknown unknown unknown unknown unknown unknown unknown