Repository navigation
Fix silent floor terminal exits and orphan worker processes - #8060
Merged
Merged
Conversation
Post-walk compile_clean and receipt-write failures now populate located refusal details, emit falsifier classification, and journal walk-terminal rows before fast-exit; the coordinator replays worker terminal detail on failure and arms PR_SET_PDEATHSIG so timed-out steps do not leave orphaned claim_executors. Co-authored-by: Cursor <cursoragent@cursor.com>
Workers that fail before floor_terminal_fast_exit now emit the same walk-terminal journal/stderr row as post-walk failures; drop duplicate coordinator-observation journal on worker failure replay. Co-authored-by: Cursor <cursoragent@cursor.com>
prctl is not available on macOS; #[cfg(unix)] was too broad for the new pre_exec arm. Co-authored-by: Cursor <cursoragent@cursor.com>
briansrls
added a commit
that referenced
this pull request
Aug 9, 2026
* Add Class B live specimen for trim pool-membership coincidence. trim is outside the substrate free-call builtin registry; it compiles only when std.algebra is already in the pool while the consumer imports std.types alone. v1-compiler-tests exercises narrow-pool failure, coincidence compile, direct-import check, perturbation stability, and FreeMonoid receiver refusal. Co-authored-by: Cursor <cursoragent@cursor.com> * Land trim free-function authority and fix TopologyEdge modeling violation. Declare importable fn trim in std.algebra with explicit trim imports across free-call sites, free_call.trim runtime dispatch, and specimen tests proving narrow-pool binding via ListedImport rather than pool coincidence. Rename TopologyEdge.port to connector_label to close bare-primitive-nicknames-concept on the fixture types overlay. Co-authored-by: Cursor <cursoragent@cursor.com> * Replace narrow-pool types overlay with minimal trim stub. Drop the verbatim std.types copy that re-minted Duration outside std.measure; the narrow-pool fixture now shadows only String/Bool kernel imports while std.algebra enters via explicit trim import from dag/std. Co-authored-by: Cursor <cursoragent@cursor.com> * Clarify trim fn body as typecheck-only identity stub. Replace the dead-branch tautology with a plain identity body and note that free_call.trim is the semantic runtime authority (review 50610). Co-authored-by: Cursor <cursoragent@cursor.com> * Fix silent floor terminal exits and orphan worker processes (#8060) * Fix silent floor terminal exits and orphan worker processes. Post-walk compile_clean and receipt-write failures now populate located refusal details, emit falsifier classification, and journal walk-terminal rows before fast-exit; the coordinator replays worker terminal detail on failure and arms PR_SET_PDEATHSIG so timed-out steps do not leave orphaned claim_executors. Co-authored-by: Cursor <cursoragent@cursor.com> * Journal pre-walk worker terminal refusals on the main return path. Workers that fail before floor_terminal_fast_exit now emit the same walk-terminal journal/stderr row as post-walk failures; drop duplicate coordinator-observation journal on worker failure replay. Co-authored-by: Cursor <cursoragent@cursor.com> * Gate PR_SET_PDEATHSIG worker spawn hook on Linux only. prctl is not available on macOS; #[cfg(unix)] was too broad for the new pre_exec arm. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com> * Fix trim call shape in floor_discovery_producer. The std.algebra trim free function declares parameter `s`, not `seg`. Co-authored-by: Cursor <cursoragent@cursor.com> * Restore coincidence narrow-pool negative control and regen trim exports. Add trim_free_call_fails_in_narrow_pool_without_algebra_coincidence compiling coincidence_specimen.dag against the two-root narrow pool (no dag/std) and assert trim does not resolve via pool coincidence; pair with a PoolCoincidence positive binding control. Regen stage0 so extdeps.uri trim import and std.algebra trim export match the emitter. Co-authored-by: Cursor <cursoragent@cursor.com> * Add HAND-RUST scaffold deferral to class_b_trim_specimen_test. Documents the import-strip Class B lane, why pool-overlay probes stay in v1-compiler-tests, and the closure-independent-binding dissolution trigger. Co-authored-by: Cursor <cursoragent@cursor.com> * decl_facts: explicit-import resolution + resolved parent/arm projection prerequisite (#7924) * Cut exact-initializer-identity successor from main (#7855 operator verdict). Lift foundation only from session/tidy-boar-761: two-root duplicate_qn fixture, fixture-scoped pool_roots, type-env projection marshal (WIP — structural blockers from operator review remain), dimensionless Rust controls, decl_facts-vs-compile population divergence note. Copy #7796 prereq bank: decl_facts_skeleton, qualified-name resolution fixtures, 12 skeleton witness cases, constructor-lexeme negative boundary tests. Does not include tidy-boar CI retry commits or unrelated branch surface. Successor scope: ExactDeclarationIdentity carrier, binding_kind gate, eight executing controls, exact whole-tree marshal/refusal census. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix pool_roots fork: one authority in fixture witness_support. The lift merged two lineages that both declared decl_facts_reflection_fixture_pool_roots with different populations. Consolidate the six-root walk in test.fixture.decl_facts_reflection.witness_support; projection witnesses import that row. Skeleton lexeme witnesses use an explicit narrower decl_facts_skeleton_fixture_pool_roots. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix CI build: sync stage0 with main gate and witness-cost surfaces. The tidy-boar lift left cli_run and v1_interpreter behind main: missing CompilerDiagnostic histogram arms, gate failure-detail builtins, and the WitnessRowCost struct claim_executor expects. Restore those surfaces without changing the decl-facts lift. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix CI build: sync stage0 with main gate and witness-cost surfaces. The tidy-boar lift left cli_run and v1_interpreter behind main: missing CompilerDiagnostic histogram arms, gate failure-detail builtins, and the WitnessRowCost struct claim_executor expects. Restore those surfaces without changing the decl-facts lift. Co-authored-by: Cursor <cursoragent@cursor.com> * Ground exact variant identity on parent/arm declaration carriers and VariantValueBinding gate. ResolvedVariantIdentity now carries full ExactDeclarationIdentity for parent and arm instead of lossy qualified-name pairs; marshaling projects parent_type and arm alongside legacy parent_qualified_name/variant_name fields. Variant-value resolution requires infer-stamped VariantValueBinding and resolves parent coproduct through the binding's parent_enum authority rather than spelling plus annotation heuristics. Co-authored-by: Cursor <cursoragent@cursor.com> * Add VariantValueBinding and module-order controls for exact initializer identity. Control 1: planted scaffold_with_data_ref specimen plus executing .dag and Rust witnesses prove a data-item reference with coproduct annotation marshals NotVariantValueProjection, not a variant value. Control 4: reversed source-file order leaves constructor parent identity unchanged. Also gate call_env_depth witness behind an armed atomic (default off) and fix namespace_alias_decl_test module gating. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix compile-clean gate: restore ensure_is_converged imports from main. Merge carry dropped gunbc.build_cache_ensure and gunbc.compile_pool_ensure imports in their witness tests. Also land control 5: ExactDeclarationIdentity lookup grain with executing witnesses for duplicate-qn distinctness and duplicate-exact-identity ambiguity. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix batch-3 decl_facts witnesses without growing migration debt. Group B: projection helpers read initializer roots (plain record + coproduct). Group A: pool-corpus duplicate bare-type index lets decl_facts marshal ambiguous variant values when witness ctx.modules is narrower than the fixture pool. Group C: delete redundant skeleton Rust test; retain only source-order seam test with a typed retirement row (no baseline bump). Co-authored-by: Cursor <cursoragent@cursor.com> * Register decl_facts_marshal_bridge in stage0 crate layout for regen. Hand-added pub mod without frontier registration made regen_verify fail: fresh emit omitted the module while the committed lib.rs carried it. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix batch-3 discovery re-eval of duplicate OtherTwo side-effect row. The entry module duplicated witness_support's ambiguous_shared_b closure loader; corpus re-eval of that local data row ran without OtherTwo in scope. Closure loading stays on witness_support's enrolled side-effect row. Co-authored-by: Cursor <cursoragent@cursor.com> * Restore ambiguous_shared_b import without duplicate side-effect data row. Import-only closure load keeps Group A green; removing the local OtherTwo data row avoids batch-3 discovery re-eval undefined-variable failure. Co-authored-by: Cursor <cursoragent@cursor.com> * Align nullary reflection witnesses with initializer projection trees. Skeleton lexeme walks on fact.node no longer apply after DeclFact.node became projection roots; assert resolved variant identity via projection helpers instead. Co-authored-by: Cursor <cursoragent@cursor.com> * Remove re-evaluable OtherTwo side-effect row from witness_support. Discovery corpus re-evaluates imported closure-loader data rows without variant imports in scope; keep ambiguous_shared_b closure load via initializer_projection import only. Co-authored-by: Cursor <cursoragent@cursor.com> * Trust importing-module TypeBinding for variant resolution; remove pool ambiguity scan. Delete cross-module bare-name candidate machinery, decl_facts_marshal_bridge, and variant_to_enum sentinel; explicit A import must resolve uniquely to ambiguous_shared_a. Update witnesses and Rust controls accordingly; remove duplicate qualified-name witness and dead closure-loader row. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix vacuous A/B witnesses and pool-grain duplicate lookup; drop Rust CI enrollment. Add ambiguous_b_specimen as a legitimate B consumer so explicit-import controls exercise both modules in the entry closure. Replace the vacuous single-module witness with discriminating positive and negative controls. Route duplicate-QN per-candidate uniqueness through PoolDeclarationIdentity instead of a parse-pool row masquerading as exact identity. Remove the enrolled Rust source-order suite from v1-compiler-tests CI. Co-authored-by: Cursor <cursoragent@cursor.com> * Rename Exact carriers to ResolvedDeclarationLocator; honest locator grain. Drop ExactDeclarationIdentity/ExactVariantIdentity aliases. Rust projection carriers are ResolvedDeclarationLocator and ResolvedVariantLocator; dag model matches. Rename duplicate-QN witness helper to pool-declaration lookup grain. Occurrence identity is not claimed anywhere on the branch. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix CI: re-home over-budget qualified-decl ref witness to long lane. cross_module_qualified_reference_emits_call_from_correct_module exceeded the 5000ms per-PR CPU budget (chronic on main, unrelated to decl_facts). Move it to test/claim/long/ with a lighter std.unicode.types fixture; keep the fast same-module control per-PR. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix node-frontier refusal: restore long witness module declaration. Changing line 1 of an existing long-lane file trips diff-before-first-declaration fail-closed in node-frontier population. Keep main's module name; only the unicode fixture and note differ from main. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix scoped v1 witness batch FLOOR-BATCH-OVER-BUDGET on CI. The scoped child exceeded its 360s batch-owned clamp (~376s measured at 93b1373 locally; CI run 31223532865 exited 1 after the same wall). Raise the v1_claim_scoped_witness_batch clamp to 480s with a receipted note (375.6s observed x 1.2 fleet margin). Initialize the scoped witness receipt header in scoped floor workers so append does not fail when the file is absent. Co-authored-by: Cursor <cursoragent@cursor.com> * Pin scoped batch clamp witness to 480s after clamp raise. witness_v1_claim_scoped_batch_is_file_grain_and_batch_owned still asserted the retired 360s batch-owned clamp; update to match v1_claim_scoped_witness_batch. Co-authored-by: Cursor <cursoragent@cursor.com> * Bind spark-standup-program-accounting doc to gunbc doc graph roots. Fleet-blocking main red: doc_graph_has_no_orphan_docs failed because docs/plans/spark-standup-program-accounting.md landed in #7972 with no HandAuthoredDocBind row. Mirrors owned-ci-control-plane-design binding. Co-authored-by: Cursor <cursoragent@cursor.com> * Revert "Bind spark-standup-program-accounting doc to gunbc doc graph roots." This reverts commit d303674. * Bind spark-standup-program-accounting.md into the doc graph (main red, from #7972) #7972 landed docs/plans/spark-standup-program-accounting.md with no doc-graph binding, so doc_graph_has_no_orphan_docs reds on main and every branch merging main inherits it. My PR, my orphan. Verified rather than assumed. Subject: 0 bindings corpus-wide. Positive control: owned-ci-control-plane-design.md, added by a DIFFERENT PR in the SAME window, returns 4 — same query, same window, one bound and one not, so the zero is a real negative and not a broken grep. Discriminating control on the fix itself: the gate returns false with the row removed and true with it restored, so the row is what closes it rather than something else in the window. Every cited symbol grep-verified before authoring — fleet_subsumption_manual_gaps_plan, dgx_spark_arrival_standing, dgx_spark_router_bindings all exist. Two plausible names I first reached for did not, and are not in the row. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Revert "Bind spark-standup-program-accounting.md into the doc graph (main red, from #7972)" This reverts commit d47618b. * Update floor batch clamp witnesses for 480s scoped-batch overhead. Main's authority witness expected 360s; this branch receipted raise to 480s in gunbc.ci_layer_roots v1_claim_scoped_witness_batch. Co-authored-by: Cursor <cursoragent@cursor.com> * Address review 50585: restore fail-closed call binding and cache keepalives. - Re-instate duplicate named/positional argument refusal (CallContractMismatch) - Route observed_peak_resident_bytes through cli_run::peak_rss_vhwm_bytes - Restore pointer-cache keepalives for param_name, var_sym, call_func_name - Align decl_facts_reflection witness note with landed nullary-value controls Co-authored-by: Cursor <cursoragent@cursor.com> * Address review 50590: honest fn-node lookup and marshal entry gate. Rename lookup_typed_item to lookup_fn_node; marshal DataItem projections from the typechecked item node and refuse when registry knows a data item but fn_nodes lacks the subject. Add dissolve-on notes for skeleton lexeme aliases and data_initializer_identity seed-retained scaffold. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix record-literal coproduct resolution to honor explicit imports. Remove the module-pool first-pick in coproduct_type_item_with_variant_children and pass the import-resolved type_item directly into marshal_coproduct_record_projection, so duplicate bare coproduct names cannot override the importing module binding. Co-authored-by: Cursor <cursoragent@cursor.com> * Route eval_decl_facts DataItem marshal through lookup_fn_node seam. eval_decl_facts now delegates DataItem node marshaling to marshal_data_initializer_projection so enrolled .dag witnesses exercise the same typechecked path as Rust seam tests. Update gap notes to match. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix review 50602: witness import and wire retained Rust test module. Import decl_facts_reflection_nullary_value_projection in the initializer projection witness module and register decl_facts_dimensionless_projection_test in v1-compiler-tests lib.rs so the retirement row matches executing coverage. Co-authored-by: Cursor <cursoragent@cursor.com> * Restore symbol_index_fill_overlay_direction_test module enrollment. Re-add the lib.rs mod line dropped during decl_facts test wiring so the fill-overlay direction regression control cited in 04_infer remains executed. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Brian Searls <briansearls1@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * Recenter the roadmap on compute, CI ownership, and fleet convergence (#8034) * Roadmap: declare the infrastructure lanes, and make main the fleet's desired state The roadmap carried no node for most of what is now the priority. Verified against the authority before writing anything: zero hits for TasksMax, sccache, compile pool, build cache, Spark, ctrl-build, or the generated-file merge driver; one incidental hit for GitHub Actions ownership. The three "watchdog" hits are observation-collapse-watchdog-restatement, a display concept that happens to share the word with the runner recovery timer. The fleet lane's six rows all say the same shape - given a stated setting, apply it and read it back. None binds MAIN as where that stated setting comes from, which is exactly the gap: the mechanism half was modelled and the authority half never was. Meanwhile the work exists in design documents nothing on the roadmap points at. generated-file-conflict-policy.md is operator-ruled with four chartered lanes and no nodes. fleet-self-converge-enforcement-design.md names the self-converge timer as missing on every host and calls it the highest-risk gap. It also cited roadmap node 2-periodic-actuation, deleted in the 2026-07-27 refresh - repointed here to its successor fleet-anti-entropy-hygiene rather than left dangling. Sixteen nodes across five lanes, three of them new: fleet +5 main-revision-authority, atomic-convergence-verdict, runner-host-convergence, runner-broker-recovery, spark-inference-serving ci-placement +1 compile-pool-envelope ci-control (new) 4 owned-execution, executed-coverage-receipt, check-projection, actions-runner-retirement, remote-build-containment generated-artifact 2 projection-registry-containment, commit-policy-census ci-cost +2 arc-reconciliation, build-once-per-subject judgment (new) 1 mechanical-review-service Focus moves from the v1 exit and guarantee-ladder lanes to the infrastructure lanes. Nothing is deleted or parked: 98 hidden rows are counted on the page and one row restores the full view. The judgment lane is deliberately off the page while its prerequisite is on it. One witness repair, and it is not cosmetic. witness_projection_is_active_only rendered through roadmap_authority(), which applies the focus, while both its negative controls name ACCEPTED nodes. Any focus that stops selecting the namespace and P-derive lanes therefore makes them absent because HIDDEN, and the claim greens while proving nothing about acceptance. It now reads the focus-independent view, where absence can only mean accepted - the same reason witness_rendered_nodes_are_declared_or_derived already reads both documents. Proven discriminating by planting an active node's headline as a negative control (FAIL) and restoring it (PASS). Executed: generated-artifact drift gate green after regeneration; 44/44 roadmap_authority witnesses; 39/39 roadmap_page; 13/13 roadmap_frontier; 7/7 roadmap_emit. No acceptance receipts recorded. 123 nodes are active and only 9 carry a bound closing check; merged is not accepted, and manufacturing receipts for 51 merges would be the rung inflation this authority spends paragraphs forbidding. roadmap-receipt-continuity is the one mechanically decidable candidate and is left for the operator. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * Compute fabric above CI, and the two nodes the wind-down inventory named Two corrections to the previous commit. ONE: the compute fabric was missing, and owning CI was standing where it should have been. Building, checking, regenerating, changing machines, serving models and eventually judging changes are four consumers of one fabric, not four execution systems, and the previous cut put the build story underneath ci-control - which makes owned CI the definition of a build rather than one caller of it. compute-exact-work-contract exact subject in, one of five typed endings out. A branch name or a working directory is not a subject. A provider that cannot serve the requested operation class refuses BEFORE running rather than answering a smaller question. compute-artifact-return-and-materialization a write-producing computation cannot report success while its declared outputs are unreachable. That is the exact live defect: a remote run finished successfully, produced nothing locally, and the stale binary left behind was compared against itself. It grounds on docs/plans/execution-spine-design.md rather than minting a parallel concept. That document is operator-SIGNED (FLAGS A-E, 2026-07-09) and its thesis is already that realization and materialization are the only downstream readers of the dependency view - which is the fabric being asked for. Minting a second scheduler beside it would have been the nicknaming DESIGN section 3 forbids, in the place it costs most. ci-owned-execution, ci-remote-build-containment, fleet-runner-host-convergence and fleet-spark-inference-serving now depend on it. ci-remote-build-containment is restated as what it actually is: a MIGRATION, whose only permitted additions are refusals, and which is deleted once its useful behaviour is a provider behind the contract. Remote execution does not live there. TWO: two nodes the inventory named that genuinely had no home. ci-floor-discovery-snapshot the ordinary and scoped workers each walk the corpus in separate processes; the second walk measures around 284 seconds. It carries the complete typed result, never a pass-or-fail flag, and a consumer that finds it absent, damaged or wrong-subject refuses instead of recomputing. Distinct from phased-single-process-ci, which removes the cross-PHASE duplicate; this removes the cross-WORKER one. compiler-declaration-floor a value was observed flowing through a field declared as the wrong type while all fifty-two behaviour witnesses over it stayed green. Five planted defects refused separately, plus an unchanged-behaviour control so the wall cannot be satisfied by refusing more of the language. No rung asserted - the claims carrier says where it stands. Focus adds compute. 100 hidden rows counted on the page. The focus note now states plainly what a focus is NOT. Off the page sits real retained work with real remaining boundaries, and nothing currently records why a hidden lane is paused or what restarts it - a lane frozen behind infrastructure, one draining to merge, and one parked because its hypothesis was falsified are three states that all read identically as absent. That is a missing dimension on the node, and it is named as landing next rather than papered over here. Executed: drift gate green after regeneration; roadmap_authority 44/44. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * Deduplication before consumers, and the mechanism chain that removes preparation The compute contract as landed said exact subject in, typed ending out. It did not say identical computations collapse to one producer, and it did not say the fabric owns admission. Without both, handing agents a compute interface is a tidier way for ten callers to launch ten cold builds of one tree - the exact failure the lane exists to prevent. compute-deduplication-and-admission two requests naming the same computation are one piece of work: one producer, everyone else a consumer of its result. The fabric decides how many cold builds a machine carries, what the pool's total demand may be, which host serves a request, and what order competing work is served in - and tells a caller which of those it waits on. Requests differing anywhere in the identity must NOT collapse, and a capacity refusal is counted rather than becoming an invisible queue. It stands IN FRONT OF ci-owned-execution in the graph rather than beside it. Owning CI adds a consumer, and a consumer added before the duplication is removed multiplies the load instead of sharing it. The mechanism staircase, each removing a DIFFERENT duplicate: ci-scoped-worker-shared-substrate the second initialised world. Isolation keeps meaning separate mutable scratch and lifetime, never recomputing facts already fixed and immutable. ci-selection-before-preparation preparation that precedes selection. Selection is sharp about meaning and blunt about cost: it concludes the corpus is irrelevant after discovering, naming, rostering and indexing it. ci-streaming-realization the barrier between preparing and executing. Its acceptance property is that the first witness executes before the last selected entry is prepared. Chained by dependency edges rather than declared as a set, so at most the next one is startable and the limit on concurrent performance mechanisms falls out of the graph instead of being prose nothing enforces. phased-single-process-ci now sits behind the scoped-substrate row for the same reason. Width two is deliberately absent from the chain and stays parked: the tested implementation shared typed bytes while each worker still built its own world. compute_consumer_admission_sequencing_note records the ruling and, separately, that native realization and shared preparation are ONE programme. Emitting a native bundle is the right destination and the miniature is decisive at its scale, but the cited production enrolment delivered zero of three native and three of three fallback, so the required path took no benefit - and native bodies surrounded by duplicated preparation would still be bad CI. Executed: drift gate green after regeneration; roadmap_authority 44/44. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * Review response: native cutover unblocked, thresholds out of RED controls, dedup bound to a real build Four review findings acted on. The lifecycle finding is answered by sequencing rather than by content and is stated at the end. NATIVE CUTOVER WAS STALE AND ORDERED BACKWARDS. native-selected-witness-bundle read "merge #7599 once controls clear" and "fallback counts trending to zero". #7599 is MERGED, and the cited production enrolment ran none of its three selected members natively and fell back on all three. So the mechanism exists, the required path takes no benefit, and the node was still describing the mechanism as the work. It is now a REPLACEMENT: emit the bundle, execute the selected population by direct call, and DELETE that population's interpreter scheduling in the same change. Interpretation survives only as a named differential control on a cadence, never as a success arm the production path falls through to. Fallback, interpreted and unavailable all at zero is the bar, not a residue to trend. Its two parent edges are removed and one is REVERSED: five-minute-ci-gate now depends on the cutover. The gate is an aggregate outcome, so making the cutover wait for it meant the one step that removes the interpreter from the required path could not start until the programme it contributes to had succeeded. The warm-merge edge went with it - no input dependency was ever shown; the bundle needs a selected population, an emitter and a toolchain, none of which merge admission supplies. The node now has no parents and is startable. witness_five_minute_ci_gate_program_chain_is_explicit CAUGHT THIS, which is what it is for. It pinned the old edge shape, so the reordering had to be deliberate rather than incidental. Updated to require the reversed edge and to assert BOTH old edges absent, so the previous direction cannot return silently. TIME THRESHOLD REMOVED FROM A SEMANTIC RED CONTROL. ci-scoped-worker-shared- substrate required "the scoped segment falls by at least three minutes". That is a tree-measured number standing in for a structural claim - the same shape DESIGN section 5 rejects for census pins. Replaced with what the row actually means: no second source loading, no second index construction, no second initialised world, same population, same outcomes, bounded scratch, no path back to cold construction. Wall, CPU and memory sit beside it as observations. A timer moving cannot claim a duplicate was removed, and a duplicate genuinely removed is not disqualified by a noisy host. DEDUPLICATION IS BOUND TO A REAL ARTEFACT BUILD. Its first slice was open to being demonstrated on two read-only checks collapsing into one verdict - the one case where duplicate work costs nothing, while the case that swamps the fleet is two cold builds. It now names two agent sessions requesting the same artefact-producing build, one producer, one compilation, one returned output, two attached consumers - and it depends on artifact return rather than standing beside it. LIFECYCLE: not in this PR, by agreement with the review. ProgramDisposition and the work-item agreement land first as their own change and this stacks behind them; RoadmapNode is constructed in 19 files and that shape change deserves an isolated review rather than riding a 24-node expansion. Executed: drift gate green after regeneration; roadmap_authority 44/44 with the chain witness green on the new direction. CI failure on b47d383 was infrastructure, established two ways: the regen job died at "Setup Rust" with rustup ETXTBSY (exit 126, 10s in, before any content ran) because runner slots share one home directory; and gunbc.roadmap_authority is absent from the 112-module regen input closure, so this change cannot alter regen output. regen_stage0 --verify run locally reports divergence count 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * Lane 1 gets its node: commit-writer admission is scheduled first, not implied Review response (PR #8034 review, finding 1): the conflict-policy charter's four lanes were covered two-of-four, and the uncovered pair included lane 1 — the one the charter titles "FIRST: the live safety hole" and the one gunbc.repo_local_git_config's authority note names as the boundary a writer cannot pass. A generated-artifact lane whose first row is the population and whose absent row is the writer reads, on the page, as if the safety hole is scheduled. Now it is scheduled, as the parent of the lane-2/3 chain, matching the charter's own transaction ("prove no unmerged index entries — lane 1 predicate"). The node transcribes the charter's two refusal arms (unmerged-stage refusal; staged-blob conflict-marker-grammar refusal) and the operator's second-pass acceptance wall (complete staged-index observation, observed-not-declared classification, provenance-receipt fixture exemption, writer bindings as countable carriers). ROADMAP.md regenerated via generated_artifact_gate main_wet; all 44 roadmap witnesses green by execution with the node in place. Deferred per the same review's recommendation, recorded here: lane 4 (keyed rosters get set/map construction semantics) and the execution-spine-design.md doc-graph binding (needs a typed HandAuthoredDocBind anchor or a plan registration) stack behind the sequenced work-item PR. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * Consolidation: quarry #8033, add the deterministic confidence lane, reframe owned CI Operator consolidation ruling (2026-08-08): #8034 is the sole roadmap authority branch; #8033's parallel edits port here rather than merging beside it. Ported from #8033, at their exact homes rather than as new identities: - placement-compile-pool-envelope absorbs the task-vs-outer-budget distinction as typed acceptance classes (ProcessAdmissionExhausted, OuterResourceEnvelopeKilled, EffectiveLimitMismatch, CompilePoolPlacementRefused) plus the 8.8%-of-visible-limit receipt. No second placement identity. - The five false-verification incident classes map to their existing exact homes in false_verification_incident_mapping_note; the proposed umbrella node does not port (a coarse identity over mechanisms this graph already decomposes is the §3 second authority). - ci-gate-contention-independent-verdict lands as the one genuine gap, recut qualitatively: host load cannot decide a semantic merge verdict; timeout is a typed execution outcome, never assertion failure; the cutoff is never widened. New lane: confidence-semantic-impact-query (owner confidence, on the focused page) — the deterministic repository-inspection product, explicitly independent of LLM/Spark; judgment-mechanical-review-service now depends on it as its grounding packet. First consumer is one real corpus query usable by a person, not a fixture suite. Reframed: ci-owned-execution's old floor is quarry + shadow oracle, not the template — obligations port only on a proven disagreement; first slice is the exact-subject → bounded population → emitted bundle → owned execution → typed receipt chain with the old path shadowing. Updated: the focus note's namespace standing now carries the 2026-08-08 exact-subject re-observation (A–D established, E/F unavailable under P2a triggers, frontier 2), superseding the stale none-of-six wording. Structure: 151 unique ids, acyclic, no dangling edges, all cited paths resolve. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * chore: regenerate drifted generated artifacts (ci auto-heal) * Spark enrollment as fleet membership, and abstraction formation on the confidence chain Two operator directions (2026-08-08), landed as four graph changes. fleet-spark-host-enrollment: srv5/srv6 stop being hand-managed boxes. The endpoint rows are DERIVED as a downstream projection joining procurement identity and network allocation — re-typing the reserved address literal refuses (§3 fork), a reverse import into the network intent refuses (the measured cycle: procurement already imports it), and a second host-identity authority refuses. Enrollment carries a GB10 inference envelope, never runner-style thread capacity. Identity converge + reach ride the installed fleet key. fleet-spark-inference-serving now depends on it; runtime, model revision, auth and the collector bundle stay in the serving row, sequenced separately per the operator's "1 and 2 now". abstraction-candidate-discovery (owner confidence): descriptive grouping of subjects observationally equivalent under a NAMED lens, carrying the distinctions the quotient would erase, nearest existing authorities, and over-collapse/demand standing. Descriptive evidence only — the normative "these should share a layer" is an explicit bridge in the judgment row, per the abstraction-calculus mode-crossing rule. Deterministic: no Spark, no LLM. First slice hard-stops on an APPLIED acceptance (consumer migrated, duplicate deleted, receipts equal) so discovery cannot become an inert analysis service. REDs plant the three negative classes: keep-distinct on a load-bearing distinction, projection-missing on a downstream-join pair (the Spark endpoint session is the live receipt), demand-absent on a consumer-less candidate. judgment-mechanical-review-service broadens to the abstraction/modeling ruling vocabulary (existing-authority, likely-duplicate, candidate-new-abstraction, projection-missing, reprime-candidate, keep-distinct, over-collapse-risk, missing-consumer/actuator/readback, unknown) and gains the discovery dependency. Chain: confidence → discovery → judgment ← spark-serving; the Spark ranks and explains over exact packets, it never becomes the repository index. Structure: 153 unique ids, acyclic, no dangling edges; ROADMAP.md regenerated via main_wet; 44/44 roadmap witnesses green by execution. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * Regenerate ROADMAP.md from the merged authorities Post-merge projection: the conflict on the generated file was taken provisionally and the bytes here are main_wet's output over the merged authority state, per the generated-file policy (regenerate, never hand-resolve). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * Bind the Spark standup accounting doc: main's floor has been red since #7972 landed it orphaned docs/plans/spark-standup-program-accounting.md landed in the #7972 omega with no HandAuthoredDocBind and no inbound link, so doc_graph_has_no_orphan_docs (and doc_graph_is_clean) have correctly refused every main push since the last green at 4cdcd57 — the doc-reachability wall working as designed, fleet-wide. Attribution receipt: replaying the doc-graph reachability rule (ROADMAP.md/DESIGN.md/runbook roots + registered plans + hand binds, markdown links as edges) over last-green main, current main, and a candidate branch shows exactly one orphan appearing in the window, this doc. The bind anchors on the physical facts the doc records — gunbc.dgx_spark_procurement spark_a3ee_reservation / spark_3bd5_reservation — and its dissolution names the follow-up the doc itself declares: the Spark standup program landing as roadmap rows, findings migrating to typed carriers, then the doc registers as a plan or deletes and the bind deletes with it. Verified by execution: doc_graph_has_no_orphan_docs and doc_graph_is_clean both PASS on this tree; both FAIL on its parent. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * Floor fixes: four boundaries back inside the brief budget; ROADMAP reprojected The floor's brief-budget witness (100 words of authored boundary per ticket, gunbc.roadmap_page ticket_brief_word_budget) redded on four nodes this branch authored or lengthened: ci-owned-execution (104), commit-writer-admission (106), abstraction-candidate-discovery (114), judgment-mechanical-review- service (114). Each boundary is trimmed to <=100 tokens with no clause of the bar dropped — overflow either compressed or already carried by the node's other fields (abstraction discovery's no-model-server fact lives in its out_of_scope). Max boundary is now 99. The sibling doc-graph reds are main's breakage (the #7972 omega landed docs/plans/spark-standup-program-accounting.md orphaned; every push since last-green 4cdcd57 refused): fixed for the fleet in PR #8053 and carried here by cherry-pick so this branch's floor does not wait on that merge. Verified by execution on this tree: witness_ticket_brief_budget_holds_and_reds PASS, doc_graph_has_no_orphan_docs PASS, doc_graph_is_clean PASS, roadmap suite 44/44, regen ExitSuccess. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * Final Spark/convergence recut per executive verdict Narrow Spark enrollment to membership/profile/endpoint with honestly Unobserved cells; join lifecycle cells into the fleet convergence verdict and name the one-producer defect; spell the inference-serving internal path; make the fleet participation migration the first abstraction-candidate-discovery specimen. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * Add ci2-complete-native-cost-receipt: whole-corpus native shadow experiment as the bounded cutover's immediate successor Operator direction 2026-08-08: keep CI2-0's bounded acceptance attainable; measure emission/compilation/execution walls separately over the complete roster, then decide from the measured warm wall whether per-PR selection remains load-bearing or is deleted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * CI-0 recut: one row, one PR — fallback deleted, whole corpus classified, emitted, executed, authoritative (operator 2026-08-08) Collapses diagnosis-precursor / bounded-cutover / whole-corpus-receipt into a single state transition on native-selected-witness-bundle; deletes the ci2-complete-native-cost-receipt row added earlier today. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * chore: regenerate drifted generated artifacts (ci auto-heal) * Add roadmap-serve-emitted-realization: dissolve the interpreted serve scaffold via emit-on-demand (srv1 outage 2026-08-08) Interpreted concat clones its accumulator (quadratic); emitted concat moves (linear). Order: emit wiring first, content-hashed bodies second, concurrency last; request deadline regardless; belt GcpProjectId fix folded in. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * Reconcile #8034 per operator review: record #8054 failed acceptance on confidence-semantic-impact-query; collapse five-minute-gate parent onto the CI-0 one-PR child (cursor review 50559 §3 finding) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * Wire compiler-declaration-floor into guarantee_ladder_edges (cursor review 50566) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * Catch #8034 up to the day's rulings: CI-0 staged (3-member merge, producer successor, v2-only terminal), general-witness-body-producer row with construct census, CONFIDENCE post-rework standing, structural fork detector as abstraction-discovery first slice Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * Operator amendment: two judgment modes over one evidence substrate, semantic judgment receipt, model-selection successor, serving infra acceptance, participation-criterion detector, forward-intent refusals Incorporates worker corrections: participation over shape as the mechanical criterion; consumer-verb stringly-enum class moved into the deterministic detector; runner-vs-inference capacity control marked not-yet-built. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * Wire fleet-runner-broker-recovery into the fleet graph (cursor review 50583) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * Wire ci-executed-coverage-receipt and ci-gate-contention-independent-verdict as prerequisites of ci-check-projection (cursor review 50587) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * De-number the gate program prose: dispatch order follows graph edges (cursor review 50595 non-blocking finding) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * CI2-0 mandate recut: one node, complete v2 witness-execution cutover in #8043 — producer successor row deleted (operator mandate 2026-08-08) The census's 3-of-9,353 was circular (restated enrollment, never attempted realization); the fixture route's limits were mistaken for compiler limits. No successor PRs; canonical-pipeline wiring, per-identity semantic-kind x realization-standing from actual attempts, predecessor deletion, and the whole-population receipt all land in #8043. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff * Dirty-worktree verification as compute's first daily consumer (operator verdict 2026-08-08) Amends compute-exact-work-contract first_slice: exact dirty-tree snapshot runs all affected (or explicitly conservatively complete) v2 tests, exact receipt at most 5s warm, BaseUnstable a distinct answer; confidence optional for narrowing, never correctness. No new node. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018gTJKsR5YdkaEc2SUGN9ff --------- Co-authored-by: Brian Searls <briansearls1@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> * Fix trim lying model per review 50629; revert TopologyEdge scope creep. Replace identity trim body with a pure-dag seam and route emitted std.algebra::trim through v1_rt::trim via rust_host_string_op_fn_emit. Restore TopologyEdge.port (revert connector_label rename from trim PR scope). Co-authored-by: Cursor <cursoragent@cursor.com> * Own pool-independent trim binding repair (#8062). Bare free-call trim now requires listed import via closure_independent registry in infer_env (func sig, global_bare, ancestry, method-bridge paths); coincidence success is refused. Six specimen tests; emit_rust adds explicit trim import. Co-authored-by: Cursor <cursoragent@cursor.com> * Enroll both Sparks; replace the matrix rectangle counts with identity joins (#8056) Three witnesses asserted the derived host/phase matrix against enrollments.length() * host_standup_spine.length(). That is a count equality over a rectangle: it cannot tell a correct matrix from one that hands every host every spine step, which is exactly the shape a participation-scoped program is supposed to make impossible. DESIGN.md section 5 rules that completeness is an identity join, not a count. They are now joins derived from each enrollment's participation-selected program: - per-enrollment cell count == program_step_count(assimilation_program_for) - duplicate-free over host_phase_cell_key - matrix hosts are exactly the enrolled roster, no strays - unmodeled count cross-checked against a disposition filter over the same matrix, rather than against a rectangle - runner enrollments still contribute exactly the whole spine srv5 and srv6 are enrolled as InferenceServingParticipation, which selects exactly four shared obligations each and no runner-only one. ENROLLED IS NOT CONVERGED: all eight Spark cells are honestly unobserved, and a witness asserts that so a future producer reds here instead of silently upgrading the claim. Positive controls keep both sides non-vacuous -- runner-only obligations do exist on runner hosts, and the spine does carry gap phases. Unblocking root fix, not scope creep: gunbc.roadmap_instrument_sandbox defined fn cell(s:) and fn row(cells:) alongside gunbc.plans.md_helpers' fn cell(text:) and fn row(cells:) -- a section 3 homonym pair. Editing this witness widened its closure enough to pull both definers into the pool, after which the plan modules' explicit `cell` imports lost to pool coincidence (#6985 Class B) and every witness in the file failed with "calling 'cell': no parameter named 'text'". The HTML-fragment builders are renamed fragment_cell / fragment_row; md_helpers keeps the plain names. The sandbox keystone still passes. Green by execution, eight witnesses, one at a time on the live tree. Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * Register service-op String wire projection method fork class (#8062). Names the infer-layer FaithfulFreeMonoid vs String receiver split, pins trim_method_form_fails_on_freemonoid_receiver, and records dissolve-on as type-node unification — separate from the trim binding-bridge repair. Co-authored-by: Cursor <cursoragent@cursor.com> * chore: regenerate drifted generated artifacts (ci auto-heal) --------- Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: gunbai-bot[bot] <289086189+gunbai-bot[bot]@users.noreply.github.com> Co-authored-by: Brian Searls <11205878+briansrls@users.noreply.github.com> Co-authored-by: Brian Searls <briansearls1@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
failure_details, emit falsifier classification, and journalwalk-terminalrows on stderr beforefloor_terminal_fast_exit.coordinator-terminalrefusals and replay the worker's located terminal detail instead of a bare "did not complete" message.PR_SET_PDEATHSIGon floor worker spawn so a foreign step timeout does not leave orphanedclaim_executorprocesses (signature 2).Test plan
cargo test -p v1-compiler --bin claim_executor walk_terminalcargo test -p v1-compiler --bin claim_executor push_ordinary_receiptMade with Cursor