Repository navigation
Get integration PR #12800 (mtcollins1 boot) green - #12883
gunbai-bot[bot] wants to merge 125 commits into
Conversation
The build job's admission step carried if_condition: none, so every mode's dispatch paid a full gunbc run to reach ResetDispatchNotThisMode. It now carries fleet_converge_host_reset_return_step_if, derived from the one mode roster. The route witness asserts the gate. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…t job's gate Drift between the admission step's mode condition and the consuming host-reset-return job's condition would skip the observer refusal in the one mode it applies to. Both derive from the roster; the witness now asserts they stay equal. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…h-landed, so the diff reduces to the drift witness Conflicts resolved per region: the workflow yml takes main's side (#12419's gate is already on main); the witness keeps this branch's stronger step-gate == job-gate assertion. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…model Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ote host, wall clock models Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…, remote request, coreutils formats, parent rule) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…tore), ipmitool observed-output rows, SOL collector/process table, SDR dump, SMpro, uptime, host console Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… the real entry (deadline refusal, pinned SOL-teardown defect, held-unit contender) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… each case's own cause; media withdrawal and SOL-drop events Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…UnimportedBareProvider on 'grant') Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…; wall clock on std.measure carriers; seed-growth row covers the file arm - Per eager-owl-205's ruling (msg_83af891b): the eleven matrix cases over the new-witness eval-step budget are typed cost-debt admissions (floor_cost_debt_admission mtcollins1_boot_matrix_typed_admissions, reason not reading) and members of a declared 4b(3) drop (gunbc.rung_drop mtcollins1_boot_matrix_new_witness_eval_step_cost, list floor_eval_step_cost_drop_boot_matrix_rows) whose restoration trigger is the natively emitted evaluation frame on the merge path. The two cases under the per-subject line are neither (a row there is stale). - Review 72230: ModeledWallClock carries EpochSecs and a signed std.measure SecondDisplacement (new, beside CelsiusDelta/ArcsecondDisplacement). - Seed-growth row names file_result_of_observation and the file/argv boundary. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Ledger-Repair-Judged: docs/design-rung-drops.md Ledger-Rows-Repaired: docs/design-rung-drops.md mtcollins1_boot_matrix_new_witness_eval_step_cost Heal-Candidate-Run: 36427507942
…e host's boot (one attempt), stale cost-debt rows removed, drop population = the 8 over-budget cases, rung-drop projection regenerated The first floor run showed every typed cost-debt row stale: the matrix cases sit under the 500ms per-subject line, where such a row blocks. The honest fix is cost, not a different exemption: operation_realization_index maps bindings by identity once per frame (each of ~190 dispatches no longer scans the list), and the SOL-loss case no longer pays a baseline attempt. Measured locally the dearest case is now 220ms CPU (was 307ms on CI), under the 302ms enrolment margin. OperationBoundTwice/BindingMatch deleted (unreachable: duplicates refuse at admission). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…e current authority Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…the total form) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…othing when nothing is due Measured on the deadline case (temporary instrumentation, reverted): the modeled dispatch was ~110ms of ~243ms CPU, and handler selection ~60ms of that -- the covering_grant fold re-derived per dispatch for ~15 distinct operations. The selection reads only the operation identity and readonly flag besides frame-fixed inputs, so the slot keeps each decided selection keyed by that complete identity. The world advance short-circuits when no BMC event, media transition or console line is due. Deadline case now 214ms local (was 307/284ms on CI runs). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ause with both identities, not invalid JSON (#12533 finding 5) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ff is the attempt's named cause (#12533 finding 4) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…r recovers it only once that process is observed dead (#12533 finding 2) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ps its SOL drop) #12423 side-chat review 5342387382: bmc_advance_pending rebuilt pending from its own accumulator after each applied event, discarding events the transition had just scheduled -- a power cycle's restore arming the after-boot SOL drop lost that drop. The advance now fires the earliest due event from the world's own pending list, applies it, and repeats on the world that transition produced. New model control a_power_restore_keeps_the_sol_drop_it_schedules (ON host, cycle at 0, restore at 5, drop at 35; quiet advances to 10 and to 40); it fails on the previous fold. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…entities (finding 5 flips) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…finding 4 flips; identity renamed in its cost-drop row) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… interrupted case flips to recovery, with live and unobservable holder controls (finding 2) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…r the real boot entry Clock jumps (ModeledClockStep, a pure function of virtual time) inside the media readiness wait: a backward step reaches the production ReadinessClockIncoherent refusal; a forward step closes the window; neither makes a handoff. cd_error_code: 16 with nothing presented (the 2026-09-27 state) refuses before the handoff; with the host on, nothing is written; an error appearing with readiness refuses; a stale lane image with 16 is replaced and booted when the stop clears the code and refused after the replace when it does not (whether it clears is an unobserved firmware fact, so both arms run); a foreign presented image is not stopped. Six of the eight join the declared eval-step drop. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ot inside the UI root page fleet-converge run 36721915217 refused the boot: the observer journalled "page.evaluate: Execution context was destroyed, most likely because of a navigation" and released as login-unobserved. It loaded "/" only to borrow an origin and ran the session POST inside that document, which was replaced under it. The session POST, the services read and the session DELETE need a cookie jar, not a page, so they now go through the browser context's request client. The sessionStorage keys the viewer reads are seeded by a context init script. The root page is never loaded; the only in-page evaluations left are on viewer.html. No retry was added, and every typed refusal and journal word is unchanged. The loopback transport's root page now replaces itself on DOMContentLoaded, and a new wet control asserts the login is seen and the root is never served. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ndle (source.min.js) Operator decision msg_b3c77f62 (option A on escalation msg_13d5904d): a modeled read-only route to the firmware's UI bundle, so vendor codes such as cd_error_code are read from the vendor's source instead of escalated. - extdeps.bmc.megarac: megarac.Ui.GetServedBundle (readonly, no session, bytes as served) - gunbc.machine_intake_mtcollins1_ui_bundle_observe: mc info firmware revision first, refuse unless 0.32; fetch; sha256; report MATCH or DRIFT against the cited 2026-09-27 read (never refuses on drift); receipt on every path - fleet-converge mode mtcollins1_ui_bundle_observe: a mode row on the shared job, reusing the fan lane's credential prelude; bytes + receipt uploaded always() - witness: revision parse, the 0.32/0.33 discriminating pair, unread revision, digest match/drift Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ion of both sides Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…tem model, hostname; KVM dry observer not yet built State at the operator wind-down, NOT green: - mtcollins1_boot_wet_on_srv1 is one call into mtcollins1_boot_on_srv1_resolving_toolchain, which takes the toolchain resolution as a function called where the observer starts (eager-owl-205 ruling A); the matrix supplies Ready/NotReady/Unresolved from production values. - gunbc.filesystem_model resolves relative paths against a cwd lens; mkdir -p and realpath -e bound. - os.Hostname.ReadShort answers srv1; the #12492 KVM wet witness world() builds through the constructors. - OPEN: the Ready arm stops at gunbc.owned_process.launch LaunchOwned -- the dry KVM observer (owned process record, journal written through #12767's kvm_journal_line, triggers, stop) is not built, so 11 matrix cases fail; the NotReady/Unresolved cases and their no-launch assertions are not yet written; re-measure the drop rows after. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ed cases, #12437 claim split, regenerated rung drops (a) The dry world now answers gunbc.owned_process LaunchOwned for the KVM observer. It starts a process visible in /proc with its "<pid> <start>" record at the pid path, and writes its journal only through gunbc.machine_intake_mtcollins1_kvm_still kvm_journal_line: connection-requested, connection-open, then established with the quoted host:port, session, attempt and gen. While live, it answers each trigger file with a still, and on the stop file it journals stop-requested, session release, browser close and stopped, then exits. A line kvm_journal_line refuses is a harness fault. Still digests are supplied beside the bytes (the gunbc.remote_host_model precedent); sha256sum answers only for those bytes. All 26 matrix cases PASS under claim_batch locally; at ec1b819, without this binding, 22 returned false. (b) a_toolchain_that_is_not_ready_... / an_unresolved_toolchain_...: no Mkdir, no LaunchOwned, no power action, and the refusal names the toolchain's standing. (c) The matrix now names the pairing claim that runs runner_browser_toolchain_here_wet for real: mtcollins1_kvm_observer_protocol_wet_witness a_held_observer_is_admitted_and_its_triggered_still_is_hash_bound. (e) #12437 (test-only) made the_build_job_carries_the_reset_observer_dispatch_admission build the whole host-reset-return job to read its if. Both gates read fleet_converge_host_reset_return_step_if, so the claim is split: the step side compares its gate to that datum, and the_host_reset_return_job_is_gated_by_the_admissions_gate holds the job side. (f) docs/design-rung-drops.md regenerated by tools.generated_artifact_gate main_wet on the merged tree; it wrote no other change. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ost rows re-cited from CI, dead-band/cost-debt reclassified, seed-growth cache named (e) The workflow claim is three claims over three producers: - the_reset_observer_dispatch_admission_step_carries_its_gate_entry_and_inputs holds the step's gate, entry and dispatch inputs over the step's own producer, under budget with no drop (claim_batch: 16,102 eval steps including shared fill); - the_build_job_carries_the_reset_observer_dispatch_admission holds structural membership in the real fleet_converge_build_job_own_steps, and is the only claim that runs that producer; - the_host_reset_return_job_is_gated_by_the_admissions_gate holds the job's side of the gate. Only the membership claim is under the new declared drop gunbc.rung_drop fleet_build_job_membership_new_witness_eval_step_cost (eager-owl-205, escalation msg_00eaf0e6, option 1). Its trigger names the capability: fleet_workflow_steps constructs only the steps a consumer demands. Its measured_by cites run 36765162766 (155,333 marginal) and run 36743802719 (174,583), and records that the cost predates gunbc#12437. (d) Every matrix eval-step row now cites PR floor run 36765162766 at 6cac35e, and the two toolchain arms join that list. The dead-band rows cite that run's CPU, and the stop-clears case (473 ms) joins them. The interrupted-attempt case (523 ms, over the 500 ms line) leaves the band for a typed cost-debt admission, as the band's self-staling rule requires. v2.test.floor_enrolment_margin the_dead_band_authority_names_exactly_the_two_app_attest_claims went false when gunbc#12533 added the matrix list, and was never re-planned. It now holds the App Attest list at exactly its two claims, and the honoured identities at exactly the two declared lists. Review 73376: gunbc.modeled_operation_realization_seed_growth names the per-frame operation_handler_selection cache: its key, scope and retention. The ProcessArgvExpansion arm was already covered, and dispatch_file exists on main. docs/design-rung-drops.md regenerated by tools.generated_artifact_gate main_wet. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…s an instant Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Reverts the integration of gunbc#12795 (merge 7393c8e): the megarac.Ui.GetServedBundle operation, gunbc.machine_intake_mtcollins1_ui_bundle_observe and its witness, the MtCollins1UiBundleObserve fleet-converge mode row and its ci_spec invoke, the 2026-09-27 bundle digest row, and the fleet-converge.yml lines. Review 73415 found the mode's step is new string-concat shell under a Scaffold whose own marker names the modeled route (v2.workflow.bash_emit), which is in use today. The mode now lands only through eager-cat-463's lane, built on bash_emit nodes. Nothing else in the tree referenced it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…etween runners PR floor run 36775473983 at b907f7d (srv3) refused four matrix cases as measured over the 302 ms margin, at 303-318 ms. Run 36765162766 at 6cac35e (srv1) had admitted the same four at 276-290 ms, at identical eval_steps. This is the margin-straddle form of gunbc.recurring_failure_mode enrolment_dead_band_has_no_representable_standing, which #12533 already repaired: a dead-band row whose fast-runner reading lies in (envelope floor, margin] is EnrolmentDeadBandWithinRunnerEnvelope, not stale. Each row cites both runs. docs/design-rung-drops.md regenerated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
PR floor run 36783312538 at 8042d17 (srv3) refused it at 320 ms against the 302 ms margin; run 36765162766 (srv1) admitted it at 283 ms, at identical eval_steps. It is the same straddle form as the previous four. It was the only blocker on that run: 0 claims failed and 0 over-cost. docs/design-rung-drops.md regenerated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ite the interrupted-attempt row Brings in gunbc#12853: a typed cost-debt reading in (line / p90, line] is EnrolmentRosterGroundWithinRunnerEnvelope rather than stale. The interrupted-attempt case's typed cost-debt reason now cites all four readings: PR floor runs 36765162766 (srv1, 523 ms), 36775473983 (srv3, 582 ms) and 36783312538 (srv3, 543 ms), and merge-group run 36793281217 (srv4, 466 ms), whose stale row dequeued this change and which that arm now holds. docs/design-rung-drops.md: the generated-artifact driver refused it, as it does whenever both sides change it. Main's projection was taken and regenerated from the merged authorities by tools.generated_artifact_gate main_wet on a binary rebuilt after the merge. By set difference, no row of either side went dark. Local on the merged tree: v2.test.floor_enrolment_margin 44/44, the seed-mirror witness 2/2, the mtcollins1 boot acceptance matrix 26/26. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 85c4f893d7
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| let halves = split(s: d as String, delimiter: unit_hold_process_marker) | ||
| if count(halves) != 2 || text_at(xs: halves, i: 0) == "" { | ||
| UnitHoldOwnerUnrecognized { raw: raw } |
There was a problem hiding this comment.
Migrate legacy boot-hold owner records
When the persistent slot is still held by a boot from the previous version, its owner is encoded as mtcollins1-boot:<run_id> without an @boot= suffix. This new decoder classifies that valid legacy owner as unrecognized, after which holder_report marks it unobservable and every subsequent boot refuses indefinitely—the exact crashed-holder state this recovery path is meant to repair. Add an explicit migration or safe handling for legacy held records before requiring the new process fields.
Useful? React with 👍 / 👎.
| if (mine as String) != (pid_namespace as String) { | ||
| HolderLivenessUnobservable { cause: join(["the holder's pid namespace is ", pid_namespace as String, " and this reader's is ", mine as String, "; its pid cannot be looked up from here"], "") as NonEmptyStr } |
There was a problem hiding this comment.
Check the boot ID before rejecting a namespace mismatch
When the runner reboots while a boot hold is persisted, the new process has a different PID-namespace link target, so this branch returns HolderLivenessUnobservable before reading the stored and current boot IDs. The boot-ID mismatch is sufficient proof that every process from the holder's boot is dead, but this ordering leaves the slot permanently held after a reboot. Compare boot IDs first, and apply the namespace gate only when they match and a PID lookup is required.
Useful? React with 👍 / 👎.
Auto-opened by session-dashboard for session
fierce-cat-881.Pushing to
fc881-integadvances this PR.Worker attestation
Before flipping this PR to ready for review, confirm each item:
npm test,cargo test) and the result.Closes #Ndirective.Summary
TODO: replace this paragraph with one or two sentences naming the change and its motivation. Reviewers read this first.
Test plan