Repository navigation
Floor time: dissolve the #6848 once-per-entry fixpoint + reconcile-miss qualified-fill (~18min residual) - #7030
Conversation
|
Bugbot is not enabled for your account, so this pull request was not reviewed. Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs. |
|
CI failure at 58e1c18 (batch-2 discovery: doc_graph_has_no_orphan_docs) was unrelated to this PR's batch-4 heavy_whole_tree_resolve flip — the origin/main merge (#7023) brought in docs/probes/emitter_residual_site_map_2026-07-21.md with no inbound link. Fixed by adding a bind: provenance line at its topical home (the ^emit_representation_mismatch milestone it documents, dag/gunbc/v1_deletion_plan.dag), rather than skipping or widening the orphan check. Verified locally: doc_graph_has_no_orphan_docs / _has_no_dangling_links / _universe_is_nonempty all PASS, and the batch-4 heavy-resolve witnesses (witness_plan_serializes_heavy_resolves, witness_optin_emit_host_isolated_from_corpus, falsifier_plan_structure_holds, etc.) all still PASS. — sent from proud-bear-438 |
…e_map_2026-07-21.md My earlier fix (this branch) and #7035 (merged to main, now pulled in via auto-merge) independently added bind: provenance rows for the same doc. Per DESIGN.md §3 (single authority), remove the redundant one added here; #7035's row in dag/tools/self_host_curated_probe_cargo.dag alone satisfies doc-graph reachability (verified locally: doc_graph_has_no_orphan_docs, doc_graph_has_no_dangling_links, doc_graph_universe_is_nonempty all PASS). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
# Conflicts: # docs/plans/floor-time-namespace-walk-regression-diagnosis.md
…band) (#7106) * docs: attribute the CI floor 45-72min band by ci-job phase and batch. Log-diff receipts on fleet runs decompose workflow vs ci-job wall time and name discovery resolve, self-host gates, and #7030 effectful recovery as the dominant buckets. Co-authored-by: Cursor <cursoragent@cursor.com> * docs: CI floor redundant-work ledger, phase TSV, and ranked levers. Re-derive run 29976989996 and five comparison arms; name per-stage duplicate work (per-entry resolve walks, ingest re-scan, cheap gates after compile) and price top levers in minutes without floor behavior changes. Co-authored-by: Cursor <cursoragent@cursor.com> * docs: register ci-floor-time audit plan in doc_graph_roots. Unblocks doc_reachability pre-push gate for the measurement-only audit PR. Co-authored-by: Cursor <cursoragent@cursor.com> * WIP: CI floor time audit: attribute the 45-72min band * WIP: CI floor time audit: attribute the 45-72min band * fix(ci): drop duplicate NFR roster + merge #7114 — restore docs-only PR diff The layering batch-1 red on 652ec65 was caused by our redundant cli_run.rs NFR backfill (already landed on main as #7114), which forced a whole-tree compile-clean on .rs and a heavier batch-1 path. doc_graph_roots bind row checked: import-bearing file, module_path in string literal — no new layering edge (not 7080-class). PR diff vs main is now docs + doc_graph only. Co-authored-by: Cursor <cursoragent@cursor.com> * docs(audit): add PR7110 270m timeout extreme-tail receipt Run 29986954853 (one-file v1_deletion_plan.dag diff) killed at ci-step timeout-minutes:270 — corpus/infra-bound tail, not diff-size-driven. TSV row + plan §1.1 + lever 5 sharpen + new lever 9 (silent ceiling ride). Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Brian Searls <briansrls@gunb.ai> Co-authored-by: Cursor <cursoragent@cursor.com>
…tion (levers 2/3/8 of #7106) (#7122) * docs: attribute the CI floor 45-72min band by ci-job phase and batch. Log-diff receipts on fleet runs decompose workflow vs ci-job wall time and name discovery resolve, self-host gates, and #7030 effectful recovery as the dominant buckets. Co-authored-by: Cursor <cursoragent@cursor.com> * docs: CI floor redundant-work ledger, phase TSV, and ranked levers. Re-derive run 29976989996 and five comparison arms; name per-stage duplicate work (per-entry resolve walks, ingest re-scan, cheap gates after compile) and price top levers in minutes without floor behavior changes. Co-authored-by: Cursor <cursoragent@cursor.com> * docs: register ci-floor-time audit plan in doc_graph_roots. Unblocks doc_reachability pre-push gate for the measurement-only audit PR. Co-authored-by: Cursor <cursoragent@cursor.com> * WIP: CI floor time audit: attribute the 45-72min band * WIP: CI floor time audit: attribute the 45-72min band * fix(ci): drop duplicate NFR roster + merge #7114 — restore docs-only PR diff The layering batch-1 red on 652ec65 was caused by our redundant cli_run.rs NFR backfill (already landed on main as #7114), which forced a whole-tree compile-clean on .rs and a heavier batch-1 path. doc_graph_roots bind row checked: import-bearing file, module_path in string literal — no new layering edge (not 7080-class). PR diff vs main is now docs + doc_graph only. Co-authored-by: Cursor <cursoragent@cursor.com> * One tree, one resolve: floor stages consume the compile-clean computation (redundant-work ledger levers 2/3/8) STEP-1 model (gunbc.ci_materialization ci_floor_resolve_receipt_note, Receipt 5): the compile-clean receipt's typed store — the main-thread process_shared_index the eager install warms — is the consumable fact; later floor stages consume it, and a stage that cannot be served is counted, never a silent re-walk. Declared cold-resolve count 4 -> 3 consciously (the note's own rewire discipline); ci.yml regenerated to match. Lever 3 (executor realization): batch_unit_lane clause (c) — a resolve-group sharing an (entry, execution_mode) some batch resolves on the memo path is colocated there, so the batch-0 cheap-gate group rides the store the eager compile-clean install warmed instead of re-deriving the same closure cold on a spawned thread (where thread-local process_shared_index is invisible); the compile anchor and emit-host then consume its walk_memo context as hits. No schedule fact added or reordered (#7088 batch-0 ordering untouched). Lever 2 (transport realization): run_gunbc_claims pools ONE gunbc child per call (gunbc.Cli.RunClaims; argv grammar gunbc.cli_invoke.cli_claim_spec 'ENTRY::FUNCTION', decoder parse_pooled_claim_spec — one grammar, both directions) instead of one child per claim row: N rows over K distinct entries pay one pool build + K entry resolves against the child's per-process shared store (resolve_entry_graph), full-ledger conjunction preserved by construction. Residue (one pool build per call; overlay manifests are composed input) counted in the redundancy ledger; dissolve-on the W3 cross-process content-keyed store. Dead per-claim argv helper gunbc_claim_run_args deleted (zero consumers). Lever 8 (partition authority): CompileCleanPartitionBoundary.entry_roots = witness_layer_roots — the roster enumerates exactly the tree the whole-tree gate compiles, never witness_layer_roots.first() ('dag' only). Closes both directions of the roster-subset asymmetry: src/v2-only .dag diffs scope to their entry closures instead of falling to the no-shard-intersection whole-tree baseline (run 29976989996's shape), and a dag-rooted touch now selects affected src/v2 importers on scoped runs. Totality glob follows the boundary (both roots; join fixed for newline-trimmed shell stdout). Every fail-closed arm unchanged. Proven by execution: executor lane tests (promotion + RED control + mode keying); pooled claim-spec grammar tests; floor_fast_plan_scopes_src_v2_entries_ both_directions on the live tree; live-tree shard totality over both roots; pooled overlay run of the real_ingest leg (3/3 PASS — stub supersession through the pooled resolve path); typed-op witnesses incl. RunClaims red control; cargo fmt clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F2Rc3TWb8FFexdVQDNbb44 * ci: retrigger pull_request run (empty) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F2Rc3TWb8FFexdVQDNbb44 * Phase-0(b) admission: enroll the two RunClaims typed-op witnesses on the bin_witness_wet roster typed_witness_invocation_test.dag is discovery-excluded by pattern; every fn in it rides the explicit bin_wet roster. The two pooled-op witnesses added with gunbc.Cli.RunClaims landed excluded-but-unrostered and the floor refused loudly (WITNESS ADMISSION REFUSAL cause=UnexecutedDeferredWitness count=2, run 29993198712) — the admission invariant working as designed; these rows give them their executing consumer. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F2Rc3TWb8FFexdVQDNbb44 * Acceptance receipts: phase-attribution row + ledger flips + lever evidence (run 29995111208) Green pull_request floor on this branch, same probe format as the #7106 baseline: compile_clean 3.58 -> 2.89, cheap gates 10.15 -> 4.65, compile-gate consume 27s -> 8ms (walk-memo hit), ingest 12.15 -> 5.24, reads_real_bytes 3.30 -> 3.13, discovery unchanged by design (lever 1 out of mandate); resolves_total 3 == declared 3 (the resolve-receipt gate's own green line). Ledger rows cheap_gates_batch1 / compile_gate_consume flip to consumes-receipt; source_root_ingest_gate to necessary-first-touch (pooled), each with this run id as evidence. Batch-4 exec-corpus anomaly (52min, the interp_recorded_fixture row) is footnoted in the row's class and attributed in the PR thread — not a mandated stage; bisection in progress. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F2Rc3TWb8FFexdVQDNbb44 * Rework R1-R3: per-gate pooled claim_batch child (the section-9 cold-child class); interim gunbc-run claim grammar deleted as a duplicate; interp_recorded per-PR enrollment reverted R1 (child-pool lever, walls honored): run_gunbc_claims realizes as ONE claim_batch child per call — claim_batch's own pre-existing pooled --entry/--function grammar, one shared MultiEntryIndex per process, per-claim verdicts NAMED by its PASS/FAIL-per-function loop, no short-circuit, exit nonzero iff any failed. The gunbc run --claim flag, gunbc.Cli.RunClaims op, and the ENTRY::FUNCTION spec grammar this branch had introduced are deleted: claim_batch already owned the pooled-claims surface (one grammar, not two). Cheap-gate transports consolidate to one call per GATE (layering 7 rows -> 1 child, was 7 cold children; extdeps 5 -> 1, was 5; the former cross-call && short-circuit inside a gate is deliberately removed — every claim reports on every run, stated in the transport notes). The pooled child stays a SEPARATE process by design — never in-executor evaluation (the executor is the 16GiB-pinned process; the child dies and frees). extdeps' private roots datum dissolved into witness_layer_roots (a nickname). R2 (batch-4 disposition at the witness grain): interp_recorded_fixture's per-PR enrollment REVERTED to OfflineLocalRecipe — its ~13+ claim_batch children each cold-index the whole workspace root (2556s on run 29995111208, the dominating row of the 53.2min batch-4 wall); too heavy for the falsifier wet lane's 600s receipt budget as-is, so local-recipe with a pooled/scoped re-enrollment dissolve-on rather than an enshrined nightly refusal. Proven by execution: pooled child 10/10 PASS over 4 entries in one process (incl. the argv-shape witness pinning claim_batch_claims_argv's exact output); RED control exit 1 with the failing claim NAMED and later rows still reporting; ingest overlay leg 3/3 through the claim_batch loader (stub supersession held); whole-tree --target dag compile green; artifact drift clean; cargo fmt clean; executor lane + scope-plan test batteries green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F2Rc3TWb8FFexdVQDNbb44 * R4 re-measure: run 30009199696 phase row + ledger evidence (floor 49.5m fits the 55m cap; wet 53.2->10.82; cheap 3.08 and ingest 6.90 land as counted residue vs the 2min targets) Fresh phase-attribution row from the post-rework pull_request run 30009199696 (head 6fe8d02, all gates green, resolves_total 3 == declared 3, peak 15.0GiB cgroup-post): ci_job 99.2->61.1min, floor step 49.5min under the 55min cap; wet wall 53.20->10.82min from the interp_recorded de-enrollment (surviving pools 20+54 rows, eval 9.35min serial). The two R1 acceptance targets MISS and are recorded as counted residue in the redundancy ledger, mechanism named per row: cheap gates 4.65->3.08min (2 pooled children at ~88s/~64s — each claim_batch child pays a corpus-denominated MultiEntryIndex build regardless of roster size); source_root_ingest 5.24->6.90min, a +1.66min REGRESSION vs the interim vehicle (4 children at ~82-118s; the claim_batch child costs ~25-30s/process more than the deleted gunbc-run vehicle — loader-parity gap on top of the shared corpus-denominated index). Dissolve-on for both: the W3 cross-process content-keyed store (or claim_batch loader parity), never a silent re-widen. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F2Rc3TWb8FFexdVQDNbb44 --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Claude <noreply@anthropic.com>
Summary
Follow-on to #6999 (bisect of the #6848 namespace-walk CI-floor regression). #6999 landed an
entry-closure memo + reconcile call-order fix, but batch-4 wall time stayed pinned at
~1111–1125s (unchanged from pre-#6999) — the ~18min residual named in dashboard work item
adhoc-21c65e1a-2ff.Root cause:
SelfHostReadsRealBytesGate,SelfHostStalenessGate, andEmitHostGatedeclared
heavy_whole_tree_resolve: falseingate_runnable_profile(
src/v2/workflow/ci_floor_plan.dag).claim_executor::run_walkroutesfalse-flagged batchunits onto freshly
thread::spawn'd threads →resolve_entry_graph→process_shared_index, athread_local!MultiEntryIndexkeyed by(thread, roots). Each spawned thread gets a coldindex, so #6999's entry-closure memo and
pool_qualified_fillcache — both living on thatper-instance index — are built and discarded on every batch-4 gate run.
DagCompileCleanGate,SourceRootIngestGate, andRegenVerifyGatealready declaredheavy_whole_tree_resolve: trueand run on the shared main-thread
walk_memopath instead — these three gates were simply aninconsistent application of an already-established pattern.
Fix: flip
heavy_whole_tree_resolve: false → truefor the three gate profiles (3 lines,.dagmetadata only — no Rust changes, no resolver/binding semantics touched).floor_heavy_resolve_chain_resource_edgesalready serializes allheavy_whole_tree_resolvegates pairwise (
witness_plan_serializes_heavy_resolves), so the three newly-heavy gates landin their own batches sequentially on the main thread — sharing the process-wide
walk_memo/process_shared_indexcache instead of paying a cold walk each, without co-residingin one batch.
Diagnosis + mechanism write-up:
docs/plans/floor-time-namespace-walk-regression-diagnosis.md§6.
Out of scope: resolver semantics,
extend_sources_to_both_closure_fixpoint/pool_qualified_fillbuild logic,rc_map_insertquadratic,UnlistedImportUsesuppression,and batch-2's separate ~301s discovery-side gap (a different scheduling shape — stays open on
the work item).
Test plan
cargo build --release --bin claim_batch --bin claim_executor: green (local, sccache bypasseddue to a container resource-contention issue unrelated to this change).
claim_batch --source-root dag --source-root src/v2 --entry src/v2/test/claim/ci_floor_plan_witness_test.dag --functions <all 21 zero-arg witness/falsifier fns> --hermetic: 21/21 PASS, includingwitness_plan_serializes_heavy_resolves,witness_optin_plan_serializes_heavy_resolves,witness_optin_emit_host_isolated_from_corpus,witness_emit_host_profile_forbids_corpus_co_residence.heavy_whole_tree_resolve: falsefor thesethree gates;
runnable_excludes_corpus_co_residencegates onprofile.memory, orthogonal tothis flag.
move off the ~1111–1125s baseline toward the ~342s pre-namespace wave 1: containment-tree resolution — layered census + name-derived loader (salvage) #6848 class, per governor-mandated
measurement protocol (
gh apijob logs,claim_executor: batch Ntimestamps).receipt line (
peak_current,hard_backoffs,forced_serial,budget_exceeded) on the sameCI run against a same-host-class baseline. Capped hosts already run 14–16GB peak against a
16.1GB budget; routing 3 more gates onto the persistent shared main-thread index changes the
residency profile even though total resolve churn should drop. A material move toward the
budget is a stop-and-report, not a ship.
Closes dashboard work item
adhoc-21c65e1a-2ff.