Repository navigation
Nightly --ignored CI lane (§1 follow-on): run the EXPENSIVE-category ignored tests on a schedule so they aren't unrun forever; DESIGN-FIRST + escalate before editing gunbc ci/ci_spec machinery — surface the category-selection mechanism (only expensive runs nightly; failing/#5424 stays red-tracked, e - #5447
Closed
gunbai-bot[bot] wants to merge 19 commits into
Closed
gunbai-bot[bot] wants to merge 19 commits into
gunbai-bot[bot] wants to merge 19 commits into
Conversation
…or pivot) Per operator directive (via bright-stag): root-cause WHY each expensive test is expensive before building any lane — most causes are defects, not intrinsic cost. Deliverable = docs/plans/expensive-test-cause-table.md. Two disjoint populations: - Population A (fierce-hawk's measured 29-test set, 28-118s, currently green): all (c) v1-seed-structural in-process full-pipeline tests. Bimodal floor (nothing 1s-28s; 2-line snippet still 69s) ⇒ a fixed whole-pipeline cost redone per test ⇒ a secondary SUITE-LEVEL (b) suspicion (memoizable prelude/closure, §2) flagged to verify. - Population B (already-#[ignore]'d subprocess/network set): mostly (a)/(b) — build_stage0 cargo rebuilds + cargo-check boundary + 2 live-network — FIXABLE, drain back to per-PR; small (d)+interim-(c) residue earns the nightly lane. Nightly-lane cfg_attr design (nightly-ignored-lane.md) PARKED for that residue. Buckets attributed from source mechanism; fierce-hawk's single-threaded user≈wall numbers corroborate. HOLD all CI-gen machinery. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…annotation Parent-authorized discriminating measurement (1 vs 3 compile_multi tests, warm, single-threaded; box ~2x loaded so ratios not absolutes): - 1 test 138.10s; 3 tests 234.96s -> marginal ~48s/test (~24s deloaded), once-floor ~90s (~45s deloaded). user~=wall confirmed. Two costs stacked: 1. ~90s once-per-process floor = test-helper module_index OnceLock (whole dsl/+v2 parse_source) -> isolation artifact, amortized once in a real suite. 2. ~48s/test (~24s deloaded) GENUINE per-test residue that sums in suite; uniform across import-less and std-importing snippets -> import-INVARIANT fixed per-compile_sources cost, NOT a memoizable import-resolve. Pure-isolation-artifact hypothesis REFUTED: ~9min real per-PR for ~21 tests, so Pop-A genuinely merits cadence-decoupling and its nightly residue is non-zero. Resolve-path: harness-other in-process (compile_multi->compile_sources-> front_end_sources->resolve_modules), CACHE-BLIND - neither discovery(~547) nor SingleClaim(~505). So the §2 disk resolve-cache does NOT drain Pop-A (wrong path + wrong cost). Two cheap wins surfaced: module_index light-scan (test-infra), Pop-B build-once (queued). Pop-A bucket stands (c): floor fixable, residue seed-perf dissolve-on self-host. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
4 tasks done
… — not (c) seed Decisive measurement via the RELEASE gunbc CLI as the (b)-vs-(c) discriminator: release compile of a std-importing snippet with --source-root dsl indexes 387 modules + resolves the closure + typechecks in 0.099s (2-line no-import: 0.010s). The release seed compiler does the whole-tree-index + resolve + compile in ~0.1s — the same work the DEBUG test binary takes ~55-80s floor + ~27-44s/test for. The per-PR rust gate runs cargo test in DEBUG (ci_spec ci_rust_gate_test_command, no --release), so the unoptimized seed compiler is ~100-800x slower = the 28-118s/test isolation numbers. => Pop-A is DEBUG-BUILD AMPLIFICATION, a build-config (b), NOT (c) seed-structural / self-host. Both earlier hypotheses refuted (import-invariant kills shared-prelude; release speed kills intrinsic-(c)). Fix levers, both fixable now: (1) [profile.test.package.v1-compiler] opt-level=3 Cargo override -> optimizes only the hot compiler dep, per-test collapses toward ~0.1s, RESTORES ~21 tests to per-PR (beats nightly lane + the refuted prelude fixture); (2) module_index light-scan (part-a, handed off) removes the extra floor. Recommended next: build-config worker applies opt-level + re-measures (proof by execution). Direct cargo test --release confirm was blocked by an sccache fleet flake; the release-CLI 0.1s is the standing decisive proof. g-z tail still pending. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…/ wall −18%, verdict-identical) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ity is stern-otter P3 not by-construction (stern-otter-43) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…-PR; no expensive-selector) Operator re-scope via quick-ant-298 on the opt-level-won trigger. Drop the expensive-category selector + Pop-A conversion list; narrow to the irreducible Pop-B residue (live-network core locked: interp_recorded_fixture + anthropic_live_e2e; cargo-check-on-emitted gray-zone + ci_ set pending warm-ram #5450 build-once drain confirmation). cfg_attr mechanism + separate nightly.yml + drift gate unchanged. CI-gen diff still escalated before landing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Residue final: 2 live-network + 5 emitted-cargo (w/ dissolution trigger to drain to per-PR if an in-process emitted-Rust-wellformedness check replaces the cargo subprocess) + 5 ci_ ratchets. #5450 drains stage0-build-only (bootstrap 307/331/374/659); opt-level drains in-process (pipeline 8273/8291/10333/21). Sequencing step-1 (residue confirm) DONE; machinery diff still escalates before landing. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
… cron wiring + nightly.yml §2-horizontal generalization (no forked 2nd emitter): expected_workflow_yml(workflow) is the ONE emit authority; expected_ci_yml/expected_nightly_yml are two wrapper rows. ci_yaml_gate is ONE gate concept over (path, expected, label); CiYamlGate + nightly are two instances. Schedule trigger cron projection wired (workflow_yaml_project.dag :94 YamlNull drop -> render_cron_schedule). nightly_workflow value (Schedule 0 7 * * * + workflow_dispatch, --features ci_nightly run step); ci_nightly feature declared in v1-compiler-tests (markings held for post-#5450). Floor-enrolled NightlyYamlGate witness (nightly_yaml_serializer_witness_test.dag). Proofs by execution (parent escalation packet): 1. ci.yml BYTE-IDENTICAL post cron-fix + prelude-factoring: CiYamlGate main=ExitSuccess, git diff HEAD ci.yml empty, existing ci_yaml serializer keystone still true. 2. cron renders by execution: pre-fix YamlNull -> witness RED + schedule block has no cron (3102 bytes); post-fix -> - cron: 0 7 * * * (3124 bytes), witness GREEN (discriminating). 3. NightlyYamlGate floor-enrolled (*_test.dag test fn) GREEN on committed, RED on perturbed; nightly.yml == expected_nightly_yml() byte-exact (nightly_main=ExitSuccess). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…t-true perturb) Parent's analysis (correct): a pure .dag witness can only build a perturbation FROM expected_nightly_yml(), and !(E+suffix == E) is constant-true regardless of the comparator — so the synthetic-perturb 'teeth' were vacuous (the tautological-control trap). Replaced with witness_nightly_yaml_committed_matches_generator: the committed nightly.yml (an INDEPENDENT on-disk fact this witness did not construct) must equal the live generator — the REAL drift detection. The comparator's own teeth are the wet probe (ci_yaml_gate.nightly_drift_wet_receipt, which writes a perturbed file and requires the real gate body to flag it) + the execution proof. Pairs with the gate-receipt wet fix (8221254). Verified on the merged head: run_ci/nightly_yaml_gate=ExitSuccess (wet receipts), ci.yml + nightly.yml byte-identical, both serializer keystones true. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Contributor
Author
|
CI red here is inherited main-red, not introduced by this PR — no fix will be pushed from this branch (doing so would §3-fork an already-open fix PR). Proof:
This PR's own gates/witnesses (ci/nightly drift wet receipts, byte-identity, the nightly serializer keystone) are green by execution. Merge readiness is gated on #5453 landing on main + a 2nd distinct review provider — both external to this branch. — sent from proud-deer-709 |
This was referenced Jun 21, 2026
briansrls
added a commit
that referenced
this pull request
Jun 27, 2026
…loor) (#5855) * Revert resolve cache to opt-in (#5789 always-on default OOMs the CI floor) #5789 made the cross-process resolved-graph disk cache always-on by defaulting its directory to `temp_dir()/gunbc-rg-cache-{user}`, dropping the `GUNBC_RESOLVED_GRAPH_CACHE_DIR` opt-in gate. That turned the cache on in CI, where it is pure cost: - Both IO paths buffer a whole cache file in memory. A hit `read_to_end`s the entire verbose-JSON file (~11x the packed 18-field-Node graph, so a 272 MiB graph is a ~3 GiB read); a miss `to_vec`s the whole JSON before write. Across concurrent floor shards this OOMs the runner (measured 14.16 GiB self-RSS at width=3 vs the 8 GiB cap). - CI hit-rate is ~0: the cache lives in /tmp (absent from ci.yml's actions/cache paths, empty on each fresh runner) and the subject key folds the compiler-exe hash, so every commit colds the whole cache. So in CI it only ever writes (which also OOMs) and never reads a graph back. This re-confirms the prior cold-CI cache-enable finding (#5447) that #5789 contradicted unreferenced. Restore the #4878 opt-in: `resolved_graph_cache_root_from_env()` returns `None` when the env var is unset, and both callers re-gate on it, so with no env var the cache is fully off (lookup and write both skipped). Both sites are pure perf shortcuts (lookup falls through to recompute, write is best-effort `let _ =`), so this changes only timing, never the resolved graph. ROADMAP regenerated from roadmap_authority.dag. Re-enabling by default is gated on a streaming IO realization (binary + `deserialize_from`/`serialize_into`, no whole-file `Vec<u8>`) — follow-on. Seed-infra (the v1 caching kernel's gating); does not cement Rust into templates and is not the substrate Share work. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fmt: rustfmt the cache import line (fix fmt --check gate) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Brian Searls <briansrls@gunb.ai> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Auto-opened by session-dashboard for session
proud-deer-709.Pushing to
session/proud-deer-709advances this PR.Worker attestation
Before flipping this PR to ready for review, confirm each item:
npm test,cargo test) and the result.Closes #Ndirective.Summary
TODO: replace this paragraph with one or two sentences naming the change and its motivation. Reviewers read this first.
Test plan