diff --git a/dag/gunbc/doc_graph_roots.dag b/dag/gunbc/doc_graph_roots.dag index 97840eb24d5..8a7a9920d68 100644 --- a/dag/gunbc/doc_graph_roots.dag +++ b/dag/gunbc/doc_graph_roots.dag @@ -107,6 +107,12 @@ data hand_authored_doc_binds: List = [ work: DeclarationRef { module_path: "v2.workflow.ci_floor_plan", decl_name: "gunbc_ci_floor_batches", field: WholeDeclaration }, dissolution: HasTrigger { text: "dissolves into a registered gunbc.plan.Plan row when its P1 lands" }, }, + HandAuthoredDocBind { + home: PlanDoc, + slug: "ci-floor-time-45-72-band-attribution", + work: DeclarationRef { module_path: "gunbc.ci_materialization", decl_name: "ci_floor_declared_resolve_count", field: WholeDeclaration }, + dissolution: HasTrigger { text: "dissolves when realization_measurement_loop Phase-0 Gantt carrier supersedes prose receipts (same trigger as ci-floor-fractal-gantt)" }, + }, ] fn hand_authored_doc_graph_roots() -> List { diff --git a/docs/plans/ci-floor-time-45-72-band-attribution.md b/docs/plans/ci-floor-time-45-72-band-attribution.md new file mode 100644 index 00000000000..1a972174f8d --- /dev/null +++ b/docs/plans/ci-floor-time-45-72-band-attribution.md @@ -0,0 +1,180 @@ +# CI floor time audit — redundant-work ledger + lever ranking + +**Status:** measurement receipt, 2026-07-23 (session vivid-fox-471). **DESIGN.md + carriers remain +authority** — prose + TSV receipts only; **no floor behavior changes** in this PR. Dissolves when +`realization_measurement_loop` Phase-0 lands a durable `.dag`-native Gantt carrier. + +**Product (operator mandate):** phase attribution is the **map**; the **product** is a per-stage +**redundant-work ledger** (what each stage recomputes that an earlier stage already computed on +the same input content) plus a **ranked lever table** priced in displaced minutes. + +**Carriers (this PR):** + +- [`docs/probes/ci_floor_phase_attribution_2026-07-23.tsv`](../probes/ci_floor_phase_attribution_2026-07-23.tsv) — per-run per-phase walls +- [`docs/probes/ci_floor_redundancy_ledger_skeleton_2026-07-23.tsv`](../probes/ci_floor_redundancy_ledger_skeleton_2026-07-23.tsv) — stage × recomputes × duplicate-of +- [`docs/probes/ci_floor_lever_ranking_2026-07-23.tsv`](../probes/ci_floor_lever_ranking_2026-07-23.tsv) — ranked levers + +--- + +## 1. Band census (map only) + +| Grain | median | 45–72 min band | notes | +|---|---:|---|---| +| Workflow (build+ci+deploy) | 55m | 56% of main runs | operator-facing "~1 hour" | +| **ci job** | **44m** | 40% | **use this for floor attribution** | +| Floor step (`gunbc ci` claim_executor) | ~35–48m | — | regen excluded (~3m) | +| **Extreme tail (action ceiling)** | **270m+** | 1 receipt | PR #7110 run `29986954853` — see §1.1 | + +The **~72 min** figure in `gunbc_ci_witness_corpus_only_batches_note` is whole-tree **emit** +infeasible for pre-push — not typical ci job wall. Scoped PRs skip compile-clean emit entirely +(`compile-clean scope: skipped`). + +### 1.1 Extreme tail — run `29986954853` (PR #7110, plan-only) + +**Receipt:** PR #7110 @ `ad44c819c`, run `29986954853`, **killed at the ci-step +`timeout-minutes: 270` ceiling** (workflow wall ~281 min including setup). Diff is **one file** +(`dag/gunbc/v1_deletion_plan.dag` only) — the most trivial affected-set path; compile-clean +should be skipped. + +**Reading:** NOT diff-size-driven. Strong evidence the tail is **corpus-denominated + +infra-bound** (serial `width=1` + memory thrash on a bad-luck runner). When width latches at 1 +there is **no recovery arm** — the run rides silently to the action ceiling (DESIGN §5 +absorbing-fallback shape: corpus-denominated cost breaks the budget later, not a typed refusal +mid-flight). Phase breakdown unavailable: log rotated on re-queue at cancel; floor never emitted +final batch receipts. + +**Lever sharpen:** extends ranked lever 5 (width=1 latch) — add **fail-fast at action ceiling** +as a separate scheduling finding (270 min silent ride vs early typed refusal). + +--- + +## 2. Receipt anchor — run `29976989996` (re-derived) + +Branch `session/gentle-raven-495`, green, srv1-01, ci job **54.9 min**, floor step **~48.4 min**. +7-batch schedule (post-#7088 cheap-gate early batch). **Whole-tree compile-clean** because diff +had no shard intersection. + +| phase | wall (min) | % of floor | +|---|---:|---:| +| preamble (plan resolve + hygiene) | 1.9 | 4% | +| compile-clean receipt (whole-tree emit) | 3.6 | 7% | +| batch 1 cheap gates (3 nodes, 1 resolve-group) | **10.1** | **21%** | +| batch 2 compile gate consume | 0.5 | 1% | +| batch 3 discovery (663 entry-groups, 2206 rows) | **12.8** | **26%** | +| batch 4 wet corpora | 1.3 | 3% | +| batch 5 emit_host | 0.1 | 0% | +| batch 6 source_root_ingest (ONE node) | **12.1** | **25%** | +| batch 7 reads_real_bytes | 3.3 | 7% | + +**Top-3 = 35.0 of 48.4 min (72%):** discovery 12.8 + source_root_ingest 12.1 + cheap gates 10.1. + +Governor receipt: `budget=16GiB` (cgroup memory.high), `max_width_reached=1`, +`measured worker share=3.36GB`, `peak_current=10.1GiB`, `cross_worker_store withheld`. +Declared cold resolves: **4** (matches `ci_floor_declared_resolve_count`). + +--- + +## 3. Redundancy ledger (product) + +Each row: what the stage computes, what earlier stage already computed on the **same content**, +and redundancy class per DESIGN §2 (duplicated / unnecessary / irrelevant). + +| stage | recomputes | duplicate of | class | receipt | +|---|---|---|---|---| +| **compile-clean receipt** | whole-tree load + resolve + typecheck + emit | — (first whole-tree touch) | **necessary** | 3.6min; builds `process_shared_index` | +| **cheap gates (batch 1)** | re-resolve witness entry + scan imports/extdeps/drift | compile-clean receipt on **same** `witness_layer_roots` | **duplicated** | 10.1min **after** 3.6min compile; 3 gates parallel, same resolve-group | +| **compile gate consume** | reads receipt artifact | compile-clean receipt | **necessary** | 27s verify only | +| **discovery** | per-entry `extend_sources_to_both_closure_fixpoint` + eval | compile-clean typed cache **in principle**; **not** per-entry walk | **duplicated per-entry** | resolve serial **643s**; `reusing process_shared_index` but #6848 walk dominates | +| **source_root_ingest** | `discover_source_root_ingest` bin full tree scan | compile-clean + discovery on same roots | **duplicated** | **12.1min** one node; separate binary path | +| **reads_real_bytes** | heavy whole-tree resolve + filesystem read | prior heavy gates | **duplicated heavy resolve** | 3.3min; serial after ingest | +| **width=1 governor** | serializes all witness work | — | **irrelevant** (scheduling) | NOT proposing cap raise; index shrink / M2 lane | +| **materialization unkeyed** | 2.19M unkeyed pure calls | keyed memo path | **duplicated (identity unknown)** | unkeyed=47% of demand; ComputationIdentity lane | + +**Key finding vs "4–5× whole-tree re-ingest" hypothesis:** declared **cold resolve count = 4** +per run — NOT four independent whole-tree cold graphs. The band is **not** four full re-ingests; +it is **one** whole-tree compile + **many per-entry walks** inside the shared index (discovery +643s serial resolve on 663 groups ≈ **970ms/group**), plus **two 12-min single-node gates** that +re-touch the tree through different code paths (ingest bin, cheap-gate scans). + +Full skeleton: [`ci_floor_redundancy_ledger_skeleton_2026-07-23.tsv`](../probes/ci_floor_redundancy_ledger_skeleton_2026-07-23.tsv). + +### 3.1 Batch-1 internal (cheap gates) + +All three gates (`layering_imports`, `extdeps_external_authority`, `generated_artifact_drift`) +PASS at the **same timestamp** — one resolve-group, wall = **max** of parallel gate evals, not sum. +Dominant cost is the **shared resolve + host-effect scan** of the gate witness closure (~10min), +not one gate beating the others in serial. Per-gate split requires `GUNBC_FLOOR_GANTT=1` on a +replay (follow-up, not this audit PR). + +### 3.2 Batch-6 / source_root_ingest — why 12 min for one node? + +Evidence from run `29976989996` log: batch 6 invokes `discover_source_root_ingest` repeatedly +(shell `test -x` preamble then long-running ingest). This is a **separate release binary**, not +the compile-clean receipt path. It re-derives source-root ingest facts from the live tree — +work **not** consumed from the typed store the compile-clean receipt populated. Same pattern on +scoped main runs: batch 5 **10.9min** (`29970583893`) even when compile-clean is **skipped**. + +--- + +## 4. Quadratic hunt (partial — historical arm) + +Fit: discovery `resolve_serial_s` vs `entry_groups` (logged per run). + +| run | class | entry_groups | resolve_serial_s | ms/group | +|---|---|---:|---:|---:| +| `29763408563` | PRE-6848 | ~500† | 99 | ~200 | +| `29819122813` | POST-6848 | ~500† | 403 | ~800 | +| `29976989996` | deep-diff | **663** | **644** | **971** | +| `29970583893` | trivial-diff | **663** | **620** | **935** | + +†PRE runs lack `adaptive pool over N entry-groups` log line; groups estimated from witness count. + +**Reading:** ms/group grew **~5×** PRE→POST (#6848 bare-reference fixpoint) while group count +grew ~30% (2087→2206 witnesses). The premium is **superlinear in per-group walk cost**, not +merely corpus size growth. **Local ptrace** on the two ~12min single-node gates is **not yet +run** (this audit PR is measurement-only); candidates: `rc_map_insert`, typecheck-env inductive +duplication, s1_closure re-walk (named in mandate). + +--- + +## 5. Mandate questions — answers + +| # | question | answer | +|---|---|---| +| 1 | What dominates each duration class? | **Trivial-diff (~48m ci):** discovery (~12m) + source_root_ingest (~11m) + effectful (~7m). **Deep-diff (+6m):** adds whole-tree compile-clean (+3.6m) + cheap gates (+10m when pre-compile ordering). No 127–159m green runs in last 500 workflow samples — operator class may be falsifier/cold-control or older fleet. | +| 2 | Why 12min for source_root_ingest? | Separate `discover_source_root_ingest` binary re-scans tree; does not consume compile-clean receipt. Batch-1 gates: parallel group, ~10min shared resolve — per-gate split needs GANTT replay. | +| 3 | How many whole-tree index rebuilds? | **1** explicit whole-tree compile emit + **4** declared cold resolves — but **663 per-entry walks** inside discovery on shared index. `fe_begin` RSS climbs 9.5→15.2 GiB across discovery despite index reuse. | +| 4 | Width=1 fleet-wide on 16GiB? | **Yes on measured runs:** `max_width_reached=1`, `cross_worker_store withheld`. Worker share ~3.4GB leaves headroom on paper but governor does not grow width (width_growths=0). Recovery = per-worker index shrink / M2, **not** cap raise. | +| 5 | #6848 / #6999 claims? | **Verified:** resolve_serial 99→644s (+545s) PRE→seed; #6999 **~0%** batch-wall recovery on comparable hosts (29855080611 vs 29819122813). Discovery loads each entry once per worker at width=1 — memo hits near zero on that path. | + +--- + +## 6. Ranked levers + +See [`ci_floor_lever_ranking_2026-07-23.tsv`](../probes/ci_floor_lever_ranking_2026-07-23.tsv). Top +three by displaced minutes: + +1. **Per-entry bare-reference fixpoint** — 8–12 min (namespace §PR-5b) +2. **source_root_ingest re-walk** — 10–12 min (module-identity lane) +3. **Cheap-gate scan after whole-tree compile** — 5–10 min (#7088 ordering may shift; sleek-crane owns) + +Config-grade follow-ups (named, not landed here): `GUNBC_FLOOR_GANTT=1` on fleet for per-gate +split; ptrace on ingest + discovery for quadratic stacks. + +--- + +## 7. Reproduction + +```bash +gh run view RUN_ID --log | rg 'claim_executor: batch|PASS \[batch|compile-clean scope|adaptive pool|discovery corpus:|\[governor\] receipt|floor materialization|floor resolve count' +``` + +--- + +## 8. Provenance + +- vivid-fox-471, 2026-07-23, log-diff by execution on runs in TSV. +- Parent mandate: sharp-bee-290 msg_eae17a34 (redundancy ledger + quadratic hunt). +- Related: [floor-time-namespace-walk-regression-diagnosis.md](floor-time-namespace-walk-regression-diagnosis.md), + [floor-shared-compute-memoization.md](floor-shared-compute-memoization.md), + [v1-run-stability-throughline.md](v1-run-stability-throughline.md). diff --git a/docs/probes/ci_floor_lever_ranking_2026-07-23.tsv b/docs/probes/ci_floor_lever_ranking_2026-07-23.tsv new file mode 100644 index 00000000000..5398acdc15d --- /dev/null +++ b/docs/probes/ci_floor_lever_ranking_2026-07-23.tsv @@ -0,0 +1,10 @@ +rank lever mechanism expected_min_recovered risk owning_thread evidence +1 Per-entry bare-reference fixpoint once-per-entry (#6848 residual) each discovery entry-group pays ~140ms namespace walk even with warm process_shared_index; 663 groups × ~140ms ≈ 5–8min discovery alone 8–12 low (perf-only; semantics frozen) namespace-resolution §PR-5b floor-time-namespace-walk-regression; resolve_serial 99s→644s PRE→seed +2 source_root_ingest_gate tree re-walk single-node batch 6/5 costs 10–12min via discover_source_root_ingest bin independent of compile-clean receipt 10–12 medium (must preserve ingest semantics) module-identity vs storage 29976989996 batch-6 12.1min; 29970583893 batch-5 10.9min +3 cheap_gates batch before compile consume (pre-7088 schedule) or gate scan duplication 10min parallel resolve-group scans layering/extdeps/drift after whole-tree compile already ran 5–10 low if derived from compile receipt ci fail-fast #7088 (sleek-crane owns ordering) 29976989996 batch-1 10.1min AFTER 3.6min compile +4 Unkeyed materialization / missing ComputationIdentity 2.19M unkeyed calls per run; 47% of demand unkeyed — duplicate pure work invisible to memo unknown until keyed medium duplicate-work graph lens 29976989996 unkeyed=2193388 duplicated=71948 +5 Width=1 latch from 16GiB slot + 3.4GB worker share governor never admits width>1; cross_worker_store withheld; discovery serial; PR7110 run 29986954853 rode 270m action ceiling on ONE-FILE plan-only diff — no recovery arm 0–4 (only if index shrinks); tail unbounded without fail-fast medium 5886 projection / M2 memo seed governor receipt; 29986954853 extreme tail; NOT proposing cap raise per mandate +9 Action-ceiling silent ride (270m timeout) ci.yml timeout-minutes:270 kills without typed floor refusal; worst case on record for trivial diff unknown (scheduling policy) low ci workflow / floor disposition 29986954853 PR7110 plan-only; DESIGN §5 corpus-denominated-later shape +6 rc_map_insert quadratic (per-entry map growth) superlinear in module count during pool_qualified_fill / reconcile budget TBD until profiled low bold-crane-271 quadratic hunt — local ptrace on 12min gates pending +7 UnlistedImportUse advisory generation (~5k rows/run) typecheck constructs advisories during every resolve 1–3 TBD low namespace advisory suppression floor-time-namespace-walk §1.5 +8 Whole-tree compile-clean on deep-diff (scope widen) deep-diff triggers 3.6min emit not present on scoped PRs 3–4 on scoped PRs only tools.dag_compile_clean_scope 29976989996 whole-tree vs 29970583893 skipped diff --git a/docs/probes/ci_floor_phase_attribution_2026-07-23.tsv b/docs/probes/ci_floor_phase_attribution_2026-07-23.tsv new file mode 100644 index 00000000000..b7e7b3316e5 --- /dev/null +++ b/docs/probes/ci_floor_phase_attribution_2026-07-23.tsv @@ -0,0 +1,8 @@ +run_id branch class ci_job_min preamble_min compile_clean_min compile_scope wall_cheap_gates_min wall_compile_gate_min wall_discovery_min wall_wet_min wall_emit_host_min wall_source_root_ingest_min wall_reads_real_bytes_min witnesses entry_groups resolve_serial_s eval_serial_s resolves_total max_width peak_current_gib worker_share_gib unkeyed_calls duplicated_keys wasted_ms fe_begin_count_disc +29763408563 main PRE-6848 17.9 0.8 whole-tree or scoped 0.8 3.4 0.9 5.7 2087 99.3 14.9 3 1 unreadable 4 +29819122813 main POST-6848 42.1 0.7 whole-tree 0.7 8.4 1.2 18.5 2128 402.7 14.9 1 unreadable 4 +29976989996 session/gentle-raven-495 deep-diff green 54.9 1.89 3.58 whole-tree baseline (no shard intersection) 10.15 0.45 12.80 1.30 0.12 12.15 3.30 2206 663 643.5 18.4 4 1 10.1 3.36 2193388 71948 23407 4 +29970583893 main trivial-diff green 47.7 1.89 0 skipped (no touched paths) 0.8 12.17 1.30 7.10 10.90 2.90 2204 663 619.7 15.1 3 0 unreadable 2413946 71948 23407 4 +29967907137 main trivial-diff green 49.2 0 skipped 0.8 12.25 1.30 7.60 11.10 2.90 2204 663 623.8 15.2 3 0 unreadable +29880571548 main POST-7030 51.0 0.9 whole-tree partial 0.9 10.10 1.40 10.00 11.20 3.20 2137 478.0 15.6 4 1 unreadable +29986954853 session/sharp-bee-290-roadmap-converge PR7110 plan-only ONE-FILE timeout@270m-action-ceiling 270+ killed unknown killed-before-receipt plan-only scoped (v1_deletion_plan.dag only) unknown unknown unknown unknown unknown unknown unknown unknown unknown unknown unknown 1? unknown unknown unknown unknown unknown unknown unknown diff --git a/docs/probes/ci_floor_redundancy_ledger_skeleton_2026-07-23.tsv b/docs/probes/ci_floor_redundancy_ledger_skeleton_2026-07-23.tsv new file mode 100644 index 00000000000..507fd73b549 --- /dev/null +++ b/docs/probes/ci_floor_redundancy_ledger_skeleton_2026-07-23.tsv @@ -0,0 +1,12 @@ +stage what_this_stage_computes prior_stage_with_same_input content_identity_shared redundancy_class duplicate_of evidence_run evidence_note +prelude_plan_resolve plan entry resolve + naming-hygiene walk + output policy install none (first touch) n/a necessary — 29976989996 t+113s to governor armed; ~1.9min +compile_clean_receipt whole-tree --target dag: load all modules, resolve graph, typecheck, emit dag artifacts none (first whole-tree touch) process_shared_index roots keyed to --source-root duplicated downstream compile_clean builds typed_module cache consumed by later resolves IN PRINCIPLE; later stages still pay per-entry walk work 29976989996 3.6min emit; receipt ok=true; scope=whole-tree baseline +cheap_gates_batch1 layering_imports + extdeps_authority + generated_artifact_drift scans over corpus compile_clean_receipt SAME witness_layer_roots closure; same process_shared_index duplicated compile_clean already typechecked whole tree; gates re-resolve witness entry + scan imports/extdeps/drift 29976989996 10.1min wall (3 gates parallel in 1 resolve-group); all PASS same timestamp +compile_gate_consume dag_compile_clean_gate_passes reads receipt only compile_clean_receipt SAME receipt artifact necessary (verify only) — 29976989996 27s — consumes receipt, does not re-emit +discovery_corpus per-entry resolve: extend_sources_to_both_closure_fixpoint + pool_qualified_fill + witness eval for ~663 entry-groups compile_clean_receipt + cheap_gates SAME pool roots; reuses process_shared_index at width=1 duplicated per-entry walk whole-tree typed cache does NOT elide per-entry bare-reference fixpoint (#6848); resolve serial 643s despite index reuse 29976989996 12.8min wall; 663 groups/2206 rows; fe_begin x4 during batch; cross_worker_store withheld width=1 +wet_corpus_small 20+53 explicit witness rows (exec/bin) discovery_corpus SAME index partially duplicated second/third discovery batch on subset; resolve serial 12s+11s 29976989996 1.3min combined +emit_host emit-host MVP smoke (cargo build subset) discovery_corpus SAME index necessary (host effect) distinct host-effect work 29976989996 7s +source_root_ingest_gate discover_source_root_ingest binary: full source-root ingest scan compile_clean + discovery SAME dag+src/v2 roots duplicated 12.1min for ONE node; runs separate bin re-walking tree; not served from compile-clean receipt alone 29976989996 12.15min wall; shell invokes discover_source_root_ingest repeatedly +reads_real_bytes_gate self_host_realized_comparison filesystem read gate compile_clean + discovery + source_root_ingest heavy_whole_tree_resolve on main thread duplicated heavy resolve 3.3min; third heavy-resolve gate serial chain 29976989996 batch 7 after ingest; #7030 routes to shared index but still pays resolve walk +governor_width width=1 entire run due to 16GiB memory.high budget + 3.36GB measured worker share all stages n/a irrelevant (scheduling) forces serial discovery; cross_worker_store withheld — cannot amortize across workers 29976989996 max_width_reached=1; peak_current=10.1GiB; NOT cap-saturated (no forced_serial) +materialization_unkeyed 2.19M unkeyed pure calls in one run all keyed stages n/a duplicated (computation identity unknown) materialization receipt: unkeyed_calls=2193388 vs keyed=2413946 29976989996 duplicate-work / ComputationIdentity lane; 47.6% of keyed+unkeyed unaccounted