Skip to content

mtcollins1 boot: phase-timing instrument (first action and verified observation per phase, cold/warm) - #12427

Merged
gunbai-bot[bot] merged 13 commits into
mainfrom
session/proud-deer-288-relay
Sep 28, 2026
Merged

gunbai-bot[bot] merged 13 commits into
mainfrom
session/proud-deer-288-relay

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Phase-timing instrument for the mtcollins1 boot step (authored by proud-deer-288; relayed and pushed by eager-owl-205 because proud-deer-288's GitHub token was rejected).

Reads the clock at each phase boundary of mtcollins1_boot_actuate under the unit hold (sol_acquire, param5_before, media_attach, handoff, terminal_watch, param5_after, power_after, sol_release). Each phase gets first_action=+Nms (offset from hold acquisition) and verified_after=Nms. The start is classified cold/warm from the handoff's observed power transition (PoweredOn = cold, PowerCycled = warm). The result is rendered into the diagnostic bundle's run section. An unreadable clock prints its cause, never 0. Order, refusals and the EDAC workload are unchanged; mtcollins1_drive_handoff now returns {outcome, start}.

Purpose: the named instrument for the operator's "boot time toward 0" lane (the side chat approved the strategy on #12419, comment 5857457629). Every later wait removal is justified by before/after readings from this record.

None of the witnesses have run yet. BuildBuddy refuses claims (no cgroup memory limit, HostBudgetUnreadable), so CI's witnesses lane is the first execution. New: dag/test/claim/machine_intake/mtcollins1_boot_phase_timing_witness_test.dag. Affected: mtcollins1_boot_run, mtcollins1_boot_diagnostic_bundle, mtcollins1_unit_hold_forged_probe (drive_handoff's return type changed).

After merge: one cold and one warm mtcollins1_boot dispatch on main for the baseline.

🤖 Generated with Claude Code

Root cause and same-class failures (operator process ruling 2026-09-27)

Instance: a // annotation inside the body of mtcollins1_boot_actuate, refused structurally by parse (DESIGN §4c), and CI caught it. The root cause is the authoring loop: the authoring session couldn't run parse before the PR opened. It has no arm64 gunbc interpreter, and BuildBuddy runners expose no cgroup memory limit, so gunbc refuses with HostBudgetUnreadable. The parse wall held; it just ran first in CI. (Fixed in b49fe38. c5c851f also supplies the new timing argument in mtcollins1_boot_run_witness_test.)

Modeled in this PR (pending the follow-up patch): a clock reading that steps backward becomes PhaseClockSteppedBack and prints as REFUSED, never as a negative duration. Control: a_clock_stepped_back_refuses_rather_than_printing_negative_time. An unreadable reading was already PhaseClockUnread.

Frontiers / scope: superseded. See the evidence table, scope exclusions, and adjacent-class residue in PR comment 5858715884 (frontiers 1 and 4 closed by fixup 3; 2 and 3 restated there as scope exclusions).

🤖 Generated with Claude Code

Brian Searls and others added 3 commits September 27, 2026 17:08
…bservation per phase, cold/warm)

Reads the clock at each phase boundary of mtcollins1_boot_actuate under the unit hold,
renders offsets from hold acquisition into the diagnostic bundle's run section, and
classifies the start cold/warm from the handoff's observed power transition. Order,
refusals and the EDAC workload are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Relayed-by: eager-owl-205 (proud-deer-288's GH token was rejected)

Copy link
Copy Markdown
Contributor

Operator priority addendum: unexpected SOL loss must be reported immediately, not discovered after the boot deadline. Detailed source/log findings and acceptance obligations are on #12423, comment 5858496324. Please have eager-owl-205 bind a SOL-owner PR/head there; keep this timing PR's scope honest (it instruments phase boundaries and is not itself a live channel supervisor).

Critical additional finding at main e66a6a6: gunbc.machine_intake_sol_hold ActivateHeld redirects stdin from /dev/zero. The reviewed upstream ipmitool code reads stdin into processSolUserInput and transmits it, so this supplies NUL input rather than a passive hold. It is a concrete transport problem to investigate before attributing all disconnects to BMC firmware. Do not substitute /dev/null blindly: the client exits on stdin EOF. Verify the deployed client identity and use an owned, silent/non-EOF supervised input realization, with zero-unsolicited-input control.

In run 36333687404 job 108662087268, the logged terminal watch rereads capture every 15 seconds for eight minutes without another collector PID/cmdline check, then returns no-begin-marker timeout. Post-watch chassis/parameter/SEL reads still complete, so 'SOL lost' and 'all BMC visibility lost' are not interchangeable. Wire the typed channel-loss event to the real watch, an actual live operator notification, and the frozen attempt/bundle; preserve current known host facts without inferring boot failure or power-off. The issue specifies discriminating process/transport tests and independent bounded management observation. No new live operation or merge verdict is issued by this comment.

…of printing a negative duration

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Copy link
Copy Markdown
Contributor

Operator process ruling for the boot path is now recorded at #12423, comment 5858637192: STOP → ROOT CAUSE → strengthen modeling for the same failure class → exact-head review → proceed. Please use it in this instrument's review; timing-source approval and readiness to dispatch hardware are separate.

Adjacent class to inspect here: a plausible number being mistaken for an established duration or verified postcondition. That includes backward/forward clock changes, unreadable clocks, skipped/unreached phases, command completion versus successful readback, and inclusive phase time being counted again as exclusive overhead. The current body at c5c851f describes backward-clock handling as pending and names other limitations; keep those distinctions explicit until the source and executed controls establish the repair.

A controlled refused-route fixture can test unreached-phase and cleanup observations without waiting for another live hardware attempt. Please do not defer that discriminator solely until a real refused bundle exists. Likewise, this instrument does not establish SOL session health or replace the collector/liveness child. The next boot needs the integrated reviewed observer and applicable readiness repairs on its actual dispatched revision, not merely this PR's timing fields.

This is process/acceptance guidance after reading the current PR description, not a fresh exact-head source verdict or permission to dispatch. End the review evidence with source status, adjacent class residue, and live-attempt readiness; link them to #12423.

@gunbai-bot

gunbai-bot Bot commented Sep 27, 2026

Copy link
Copy Markdown
Contributor Author

Response to the lane requirements (#12423 comment 5858637192; the addendum on this PR). Fixup 3 (mtcollins1 boot phase timing: boot-based clock, per-phase verdicts, not-reached phases, and a controlled refused route) is being relayed on top of fe1b56d.

Class: a plausible number mistaken for an established duration or a verified postcondition. Each instance you named, and what now holds in source:

Instance State in source Control (all in mtcollins1_boot_phase_timing_witness_test)
Backward clock step Closed by construction: the clock is /proc/uptime, the kernel's boot-based clock (new extdeps.linux.proc_uptime and linux.Procfs.ReadUptime). A reading below the one before it still refuses as PhaseClockSteppedBack and is never printed as time. a_clock_that_goes_back_refuses_rather_than_printing_negative_time
Forward clock step Closed by the same change: NTP does not step the boot-based clock. The earlier frontier row is withdrawn. the_real_uptime_clock_reads_and_does_not_go_back (the real procfs route on the runner)
Unreadable clock PhaseClockUnread prints its cause, never 0. A malformed /proc/uptime refuses. an_unread_clock_prints_its_cause_not_a_number, proc_uptime_parses_centiseconds_and_refuses_malformed_text
Command completion vs verified readback Each phase that ran carries its own verdict, derived from its typed result: BootParam5Reading, the attach/handoff exits (which themselves require readback), PowerAfterAttempt, HeldLeaseObservedState. A refused phase prints REFUSED_after=… (cause) and never verified_after. a_refused_attach_marks_later_phases_not_reached_and_keeps_cleanup_times
Skipped or unreached phases PhaseNotRan prints not reached (reason) with no number. A refused SOL acquire records its own span and marks every later phase not reached. the same refused-route control, plus a_refused_sol_acquire_holds_nothing_after_it
Inclusive time counted again as exclusive The phases that ran partition hold → release (each begins at the reading that ended the one before). The SMpro power-on pass is its own phase and no longer sits inside terminal_watch. the_phases_that_ran_partition_hold_to_release_on_both_routes

Controlled route, no hardware. The spans come from a pure mtcollins1_boot_phase_spans(readings, verdicts). mtcollins1_boot_actuate is its only production caller and feeds it the readings it took under the hold. The witness drives both a completed route and an attach-refused route through it.

Scope limits (declared exclusions, not frontiers):

  • This instrument does not observe SOL session health and does not replace the collector or the liveness child. The SOL-loss requirements belong to the SOL-owner PR that eager-owl-205 binds on mtcollins1: absorb manual virtual-media recovery into the held boot path, with readiness and residue closure evidence #12423.
  • terminal_watch reports the time from verified power-on to the verified terminal observation, as its definition says. It makes no claim about how that time splits between hardware (firmware, memory training, the census workload) and poll detection latency. That split is the evidence the poll-cadence PR must supply for its own before/after, and it isn't asserted here.

Source status: fixup 3 is written against fe1b56d. It has not executed in this session: there is no local interpreter, and BuildBuddy refuses claims with HostBudgetUnreadable. CI on the relayed head is the first run.
Adjacent class residue: none known within the claimed scope above.
Live-attempt readiness: not ready and not requested. This PR is timing only. The next boot needs the integrated reviewed SOL observer and the readiness repairs on its dispatched revision (#12423).

Brian Searls and others added 4 commits September 27, 2026 18:50
…ot-reached phases, and a controlled refused route

The phase clock reads /proc/uptime (extdeps.linux.proc_uptime, a new ReadUptime on
extdeps.linux.procfs), so a wall-clock step cannot move it. Each phase that ran carries
its own verdict (verified or refused with cause), phases that never ran print as not
reached with no number, the SMpro power-on pass is its own phase, and the spans are
built by a pure mtcollins1_boot_phase_spans that the witness drives through an
attach-refused route and a partition check.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…IGN §4c); the rationale lives in extdeps.linux.proc_uptime
…gebra; read list elements with skip/first instead of the unrostered bare get; string_length for the fraction
…verdict arms directly

/proc/uptime parses straight to Millisecond at the extdeps boundary (the hundredths are a
format fact of proc_uptime(5)); phase durations are measure_sub, whose
MeasureSubtrahendExceedsMinuend is the clock-went-back refusal. The Bool helper over the
phase verdict is gone: mtcollins1_boot_phase_spans matches PhaseVerified/PhaseRefused.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 27, 2026

Copy link
Copy Markdown
Contributor Author

Review 71860 addressed in fixup 4 (being relayed on top of 2829435):

  1. Bare Int time. extdeps.linux.proc_uptime now parses %lu.%02lu directly to std.measure Millisecond (UptimeObserved { uptime: Millisecond }). The hundredths are a format fact of proc_uptime(5), declared there as proc_uptime_fraction_units_per_second; the conversion goes through milliseconds_per_second(). Phase durations are measure_sub(end, begin): MeasureDifference gives PhaseElapsed { elapsed: Millisecond }, and MeasureSubtrahendExceedsMinuend gives PhaseClockSteppedBack { begin, end }, carrying both readings. The product layer does no arithmetic by hand, and the witness checks adjacency with the instrument's own duration instead of comparing counts.
  2. Bool helper. phase_verified is deleted, and mtcollins1_boot_phase_spans matches PhaseVerified / PhaseRefused directly. I did not reuse the ObservationEstablished/ObservationRefused vocabulary from std.goal_assessment: ObservationEstablished carries an observed value, and a phase verdict has none to carry. Its value (the readback, the power state) already lives in the run record. So mapping a phase onto it would mean either inventing a unit payload or duplicating the record.

…egister answered

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

@briansrls briansrls left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Exact-head verdict: HOLD / REQUEST_CHANGES

Reviewed 02e7621, reconfirmed as the current open, non-draft, mergeable head against main base fd0b879. I inspected the eight-file diff, the exact-head actuation/phase-builder/rendering routes and witnesses, comment 5858715884, and the latest SMpro fixup. This is a source verdict, not a live-dispatch authorization.

The SMpro tightening is accepted. The actual smpro_phase_result now consumes mtcollins1_smpro_pass_result; only at least one SmproWordRead produces PhaseVerified. An empty pass and a pass made only of NAK/transport/malformed/not-sent/unencodable results do not. The added control uses smpro_probe_outcome on raw NAK, timeout and successful-word specimens, plus the empty case. This means 'at least one register value observed', NOT 'all probes succeeded' or 'hardware healthy'. The detailed outcomes remain in the existing SMpro record.

Two defects in this instrument's own stated contract block approval. Neither requires waiting for the general BMC simulator, host-binding work, or fixing SOL/media in this timing PR.

1. P2 — Refused routes do not partition the recorded interval, and the committed partition witness is false by its own data

mtcollins1_boot_phase_spans suppresses handoff/SMpro/watch spans after refused attach (and SMpro/watch after refused handoff), but always begins PhaseParam5After at r.terminal_observed.

The existing readings()/attach_refused() fixture has:

  • last reached phase before cleanup: media_attach [4000, 9000] ms;
  • next reached phase: param5_after [500000, 501000] ms.

The gap is 491000 ms. All PhaseRan durations sum to 15000 ms, whereas released - hold_acquired is 506000 ms. Therefore the refused conjunct of the_phases_that_ran_partition_hold_to_release_on_both_routes must be false under the intended list/measure semantics; the adjacent check at that boundary is not zero. I independently cross-checked these sums in Python, not by running claim_batch. The rendering witness even expects param5_after: first_action=+499000ms on this same refused route.

This is not only an unrealistic fixture: production takes fresh power_verified and terminal_observed procfs readings even when those operations were skipped. Those readings need not equal attach_verified (or power_verified on handoff refusal). Time spent in that intervening dispatch/instrumentation path then belongs to no PhaseRan interval. It may be small normally, but the partition/coverage guarantee is false.

Repair the boundary model/producer together. Either derive the next measured boundary from the last phase actually reached and deliberately account for intervening bookkeeping, or represent that real interval explicitly as orchestration/instrumentation overhead. Keep unreached hardware phases non-numeric. Do not fix this by making only the test timestamps equal, reporting skipped work as executed, discarding the interval, or deleting the partition assertion.

Add/refine tests for attach refusal, handoff refusal, terminal-watch refusal, SOL-acquisition refusal, and the completed route. Require both adjacency/coverage and sum equality with independently distinct readings across skipped branches. Exercise the actual pure builder and the boundary-selection logic the production caller uses. This does not require a hardware run. Directly execute the complete timing witness file; a green unrelated/diff-selected claim does not establish that this partition claim ran.

2. P2 — The new format boundary accepts malformed/incomplete uptime documents as observed time

proc_uptime_parse examines only words.first(). It does not establish the declared two-field, single-record format; after a valid first field it ignores everything else. Source-derived examples accepted as UptimeObserved include:

350735.47
350735.47 not-a-number
350735.47 234388.90 extra
350735.47 234388.90\n1.00 0.00

That contradicts the format authority and the witness's explicit '%lu.%02lu %lu.%02lu; anything else refuses' contract. The first field's lexical check is also integer conversion plus a two-character fraction length, not the declared digits.two-digits recognizer; signed integer spellings need explicit negative controls rather than relying on parse_int to reject them. The inspected emitted runtime's parse_int is s.parse::<i64>().ok(), which is a signed-number parser, not this procfs grammar.

Establish the declared record shape and unsigned decimal lexical fields before constructing Millisecond, preserve the offending text on refusal, and test missing/extra fields, an extra record, malformed idle field, and signed/non-digit fraction forms as well as the existing valid/negative/truncated cases. The unused idle quantity does not need a new semantic comparison (on multicore it need not be <= uptime), but malformed framing must not silently become a trustworthy complete reading. Keep conversion at this existing boundary and use the existing representability/checked-arithmetic authorities for its finite numeric realization.

Primary format grounding checked in this review: Linux v6.8 fs/proc/uptime.c calls ktime_get_boottime_ts64, applies the time-namespace offset, and prints %lu.%02lu %lu.%02lu\n; proc_uptime(5) identifies the two fields and says uptime includes suspend. Sources: https://github.com/torvalds/linux/blob/v6.8/fs/proc/uptime.c ; https://man7.org/linux/man-pages/man5/proc_uptime.5.html . This is format/clock grounding, not a claim about the deployed runner's exact kernel build.

Accepted implementation and scope

  • The named instrument is genuinely connected: actual proc_uptime_read calls feed the builder in mtcollins1_boot_actuate; the timing value passes through actuation -> attempt -> MtCollins1BootRunRecord -> the existing diagnostic-bundle renderer, including the SOL-refused route. The supplied timing argument was added to the existing run-record witnesses without replacing their original obligations.
  • The measured clock is the executing runner's boot-based clock, not the target host's printk uptime or BMC clock. Conversion to Millisecond remains at the procfs boundary; measure_sub keeps end-before-start as a separate result, and unread endpoints are not replaced with zero.
  • PhaseNotRan emits no timing number. Phase result and duration validity are separate: a successfully observed postcondition may have unread timing, and a refused operation may have a valid elapsed duration.
  • I found no intended reordering of the original BMC operations or change to the workload/refusal branches in this delta. Handoff still evaluates once; its existing observed transition supplies start classification instead of issuing another chassis read. Clock reads are new observation effects; their failure result is not used to grant a boot or relax an original refusal.
  • power_after being verified means its power state was observed, not that OFF or ON satisfies a particular boot outcome. The stricter timing disposition for inaccessible SOL release does not by itself repair or alter the pre-existing overall teardown judgment.

Adjacent-class answer and corrections to the claims

Still open within this instrument: successful-looking timing assembled from incomplete input or a reached-phase list that does not cover the actual elapsed interval. The two blockers above are concrete instances of that class. Do not claim 'none known' until they are closed and the controls execute.

Also make the measurement boundary explicit in the record/docs: hold_acquired is sampled at entry to actuation AFTER acquisition; released is sampled after SOL release and BEFORE unit_hold_release. This is not a measurement of lock acquisition/release, preparation, bundle generation, whole workflow latency, or time to the first actual network byte. first_action is currently a phase-boundary offset and includes unseparated local instrumentation/bookkeeping. Either label that honestly or add the additional observations needed for the stronger name; do not claim a complete dispatch or hold-transaction partition from these endpoints.

The normalized unit is milliseconds, but /proc/uptime's emitted resolution is 10 ms and includes suspend. Equal reads can legitimately produce 0 ms at this grain; 0 ms must not be advertised as proof of zero work. A small before/after comparison also includes the observer's process/scheduling overhead. Keep same runner/boot/time-namespace context for intervals. General cross-host/restart stitching is not a capability claimed by this local instrument.

Cold/warm here is explicitly power-up from OFF versus observed cycle from ON, not release-pack cache miss/hit, image-cache state, or a warm CPU-reset category. Preserve that evidence basis in performance reports.

The controlled attach-refused test is a pure record-builder test, not an executed BMC-orchestrator refusal; its comment is appropriately explicit. Likewise the_real_uptime_clock_reads_and_does_not_go_back receives the newly declared constant mock_response when run hermetically. That is legitimate dry parser/dispatch coverage, but only an explicitly wet local-procfs execution establishes the real-read claim. No hardware boot is needed for that local clock test.

The PR body's proposed 'after merge, one cold and one warm dispatch' does not supersede #12423's integrated live-attempt hold. Keep the exclusion explicit: this instrument is not the SOL supervisor or the media readiness repair. When those PRs are composed, retain the new media/observation states and all actual phase boundaries rather than resolve type conflicts by dropping one lane's record.

Evidence and disposition

I did not run the .dag engine, compiler/claim suite, simulated BMC, or any live hardware. The interval computation above is an independent arithmetic cross-check of the committed fixture, not a passing/failing claim_batch receipt. The reviewed body/comment disclose earlier test-execution limitations; I have not verified dashboard 71882's underlying run. At my last exact-head CI read, clippy succeeded and compiler/floor/emit-build were still running. No code, workflow, or hardware was modified by this review.

Source: HOLD. Adjacent class: OPEN at the two named timing boundaries. Live-attempt readiness: existing HOLD unchanged. Keep the clock/SMpro separation and production carry; repair the refused-route coverage and parser contract, run the complete named controls, then rebind the exact head.

]
},
[
PhaseRan { phase: PhaseParam5After, result: v.param5_after, begin: r.terminal_observed, end: r.param5_after_read },

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Preserve coverage when intermediate phases never ran

This always begins cleanup at terminal_observed, although the last PhaseRan before cleanup ends at attach_verified after attach refusal (or power_verified after handoff refusal). The committed fixture has attach_verified=9000 and terminal_observed=500000, so the refused-route partition omits 491000 ms and its own adjacency witness must fail. Production also takes distinct clock reads through the skipped branches. Account for that interval through the actual reached-route boundaries or an explicit orchestration span; do not merely make the fixture timestamps equal or assign numbers to phases that never ran.

Comment thread dag/extdeps/linux/proc_uptime.dag Outdated
// ANYTHING BUT `digits.two-digits` REFUSES with the text it read; no field is defaulted.
fn proc_uptime_parse(text: String) -> ProcUptimeRead {
let words = filter(split(s: trim(s: text), delimiter: " "), w => w != "")
match words.first() {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Validate the declared uptime record before minting an observed clock reading

Only the first space-delimited word is consumed. 350735.47, 350735.47 not-a-number, and a valid first record followed by another are all accepted, contrary to the declared two-number/single-record format and 'anything else refuses' witness contract. Validate the framing and unsigned fixed-decimal grammar, then project the first field to Millisecond; the idle quantity can remain semantically unused. Include signed fraction spellings as controls: length==2 plus a signed parse_int is not itself a digits.two-digits recognizer.

… route; strict /proc/uptime grammar

The actuation now records one step per phase in the branch that performs or skips it (a ran
step carries its verdict and the one reading taken after it; a skipped step carries no
reading), and mtcollins1_boot_phase_spans threads each begin from the previous ran step's
end, so every route -- completed, attach-, handoff-, terminal- and SOL-refused -- is
adjacent and covers hold to its last phase by construction. /proc/uptime parses only the
kernel's record: one line of exactly two unsigned digits.two-digits fields, checked on the
characters before any integer is parsed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 27, 2026

Copy link
Copy Markdown
Contributor Author

Side-chat review 5332054527 addressed in fixup 6 (being relayed on top of 02e7621):

P2-1: refused routes did not partition.

  • Root cause: the span builder and the boundary producer each decided reachability separately. The producer also took readings for operations it skipped.
  • The builder now consumes the route the actuation actually took: MtCollins1BootRoute, a fixed-shape record with one MtCollins1PhaseStep per phase.
  • mtcollins1_boot_actuate creates each step in the branch that performs or skips its phase. A step that ran carries its verdict and the one /proc/uptime reading taken right after the operation. A skipped step carries its reason and no reading, so no reading exists for an operation that was not performed.
  • mtcollins1_boot_phase_spans threads each begin from the previous ran step's end, one binding per phase. So the spans are adjacent and cover hold → last ran phase by construction, and unreached phases stay non-numeric.
  • No fixture timestamps were equalised and no assertion was deleted. The refused-route fixtures carry the times the actuation would actually read on that route.
  • Control: every_route_is_adjacent_and_covers_hold_to_its_last_phase asserts adjacency (using the instrument's own duration) and total coverage (the ran spans' sum equals hold → route end) for the completed, attach-refused, handoff-refused, terminal-refused and SOL-refused routes. There is also one rendering claim per route.

P2-2: /proc/uptime framing and grammar.

  • proc_uptime_parse accepts exactly the kernel record %lu.%02lu %lu.%02lu\n: one line ending in one newline, split on single spaces into exactly two fields.
  • Each field must be unsigned digits.two-digits. That is checked on the characters with std.decimal decimal_digits_only before any parse_int, and the idle field is validated too.
  • Every refusal keeps the text it read.
  • Controls: proc_uptime_refuses_bad_framing (empty, no newline, one field, three fields, a second record, a leading blank, a double space) and proc_uptime_refuses_malformed_and_signed_fields (non-numeric, no fraction, one or three fraction digits, empty seconds, -/+ on the seconds, a signed fraction, a signed idle field). proc_uptime_accepts_exactly_the_kernel_record is the positive control.

Source status: written against 02e7621. It has not executed anywhere yet; this session still has no interpreter, so CI on the relayed head is the first run.
Adjacent class residue: none known in the claimed scope.
Live-attempt readiness: unchanged, not requested (#12423).

gunbc-ci-auto-heal added 2 commits September 27, 2026 21:16
…bare ends_with (gunbc.rust_item_scan defines one, so the floor would refuse the unimported bare name)
…ll broke a line after 'begin:', which parse refuses)

@briansrls briansrls left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Exact-head rebind — APPROVE-MERGE (timing instrument only)

Reviewed 4d94a4e, reconfirmed as the current open, non-draft, mergeable head. I inspected the three-commit successor delta from held 02e7621 and the exact-head parser, route builder, production actuation and timing witnesses. The two concrete source findings in review 5332054527 are repaired. This approval does not clear #12423's live-attempt hold.

1. The producer and builder now agree on reached phases

The production branch that actually calls handoff records PhaseStepRan with the reading taken after that call. A skipped handoff has PhaseStepSkipped with no reading; the same separation holds for SMpro and terminal watch. The previous unconditional power_verified/terminal_observed samples on skipped branches are gone.

mtcollins1_boot_after threads the previous real endpoint through skipped steps. Therefore param5_after begins at the attach endpoint on attach refusal, at the handoff endpoint on handoff refusal, and at the terminal endpoint when the watch ran. The model explicitly assigns intervening bookkeeping to the next reached phase instead of dropping it or pretending an unexecuted hardware phase ran. SOL-acquisition refusal uses the same route vocabulary and builder.

The controls now cover completed, attach-refused, handoff-refused, terminal-refused and SOL-refused routes, requiring adjacency AND total duration. I independently cross-checked the committed fixture arithmetic in Python: total/outer intervals are respectively 506000/506000, 15000/15000, 76000/76000, 534000/534000 and 2500/2500 milliseconds. This is an arithmetic check of supplied fixtures, NOT execution of claim_batch. The producer change, not merely altered fixture timestamps, fixes the original missing-interval bug.

2. The procfs boundary now admits the declared record grammar

The parser requires a trailing newline, excludes another record, requires exactly two space-separated fields, and validates unsigned decimal characters plus exactly two fractional digits in BOTH fields before integer conversion. Missing idle, malformed idle, extra fields/records and signed fractions no longer become observed uptime. The new controls exercise those exact specimens and retain ordinary and zero-valued positive records.

The conversion remains at extdeps.linux.proc_uptime, returning std.measure Millisecond; duration subtraction remains measure_sub with a distinct backward result. The final parse fix is the route_end let-chain spelling, not a change to the clock or hardware decisions.

Preserved consumer and result semantics

The instrument remains wired through actuation -> attempt -> run record -> diagnostic bundle. Each operational result is independent of timing validity; skipped phases carry no numeric duration. The accepted SMpro fix still requires at least one SmproWordRead, not merely that probes were attempted. All-probe refusal/empty results remain refused. This establishes at least one word, not complete SMpro health.

No new hardware operation is added by the successor delta. Existing parameter-5, power and SOL ordering and same-proof release remain. Cold/warm remains the handoff's observed power-up/cycle category, not cache hit/miss.

Adjacent-class answer and scope notes

The former missing-interval and partial-record cases are closed at source. Remaining measurement limitations must not be promoted into stronger claims: the first sample is at actuation entry AFTER hold acquisition; the final sample is after SOL release BEFORE unit_hold_release. Preparation, lock-transaction time, diagnostics and whole-workflow latency are outside this record. 'first_action' is a phase-boundary offset; the later phase includes local bookkeeping and observer overhead. The instrument does not locate the first network byte or separate firmware work from poll-detection delay.

The same-runner/boot/time-namespace basis and the format's centisecond resolution remain assumptions of this local instrument. Equal samples may yield zero at that grain, not proof of zero work. Do not stitch records across runner incarnations as though they shared one clock. General full-range decimal-to-millisecond overflow coverage is also not demonstrated by these fixtures; the conversion still uses ordinary Nat arithmetic. Add a checked scaled-conversion boundary/control before claiming this is a general arbitrary-input numeric parser. I do not treat that broader claim as established by this local procfs timing approval.

The pure route matrix is not the effectful BMC acceptance harness, and the hermetic ReadUptime mock is not a wet clock measurement. Execute the complete named timing witness file and the local procfs wet control as execution evidence; do not infer those runs from unrelated green claims. No live boot is needed for either.

Landing/evidence

At my latest exact-head read, compiler and clippy succeeded; floor and emit-build were still running. This is SOURCE APPROVAL conditional on normal exact-head/merge-queue gates, not a claim that all execution checks are green. I have not independently rerun the .dag suite or verified a completed full timing-file receipt. Preserve that evidence limitation in the PR until its actual execution is recorded.

I did not regenerate, dispatch, merge or access hardware. The PR body's proposal to run cold/warm boots after merge does not authorize them: the integrated media/SOL/harness holds still apply. When composing those lanes, rederive the actual route and retain every new observation boundary rather than resolve conflicts by dropping records.

Source: APPROVE-MERGE at this SHA after normal checks. Original source HOLD: closed at the two reviewed defects. Adjacent measurement limitations: explicit above, not claims of full workflow observability. Live-attempt readiness: existing HOLD unchanged.

…eads the route end from the last span that ran

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 27, 2026

Copy link
Copy Markdown
Contributor Author

Review 71923: fixed. mtcollins1_boot_route_end is deleted. It repeated the phase threading, and its only caller was the witness. The coverage claim now takes the route end from the last PhaseRan span that mtcollins1_boot_phase_spans returns (test-local last_end over ran_bounds). So adjacency and total coverage are now checked against the production spans themselves, not a second copy of the threading. — sent from eager-owl-205

@briansrls briansrls left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Exact-head rebind — APPROVE-MERGE (timing instrument only)

4aef708. Reconfirmed current open/non-draft/mergeable head. Compare against approved 4d94a4e is exactly one successor and two files: deletion of the duplicate mtcollins1_boot_route_end chain and replacement of its test-only use. No production actuation, span construction, clock parsing, hardware order, workload, or verdict logic changes.

The witness now obtains bounds from the actual mtcollins1_boot_phase_spans result once. It still checks adjacency from the supplied hold, sums the emitted executed-span durations, and compares that sum with hold-to-last-emitted-end. Its local last_end traverses the returned pairs rather than duplicating the fixed phase-order chain. Malformed pairs cannot silently satisfy the whole predicate because adjacent_from separately rejects them. The real ran_bounds producer supplies two-element pairs.

I inspected the current surrounding controls: all five supplied route cases remain; completed and attach-refused cases still require their particular final sol_release lines; the refused-handoff and refused-terminal assertions, not-reached assertions, clock-refusal checks, and strict procfs specimens are not weakened by this commit. The production builder continues to own phase sequencing. No blocker in this delta; approval 5332382476 is rebound at this SHA.

Adjacent-class qualification

The new property establishes continuity/totality over the spans actually returned. By itself it cannot establish that a producer did not omit a final phase (and an empty output is vacuously contiguous). The explicit route/phase expectations are therefore part of the test obligation, not redundant decoration. Preserve them, and when expanding the route make the expected reached phase population and expected last reached observation explicit for each route. Do not reintroduce a second production threading function merely to serve as the test oracle. I found no production omission in this unchanged builder and am not treating this qualification as a new source blocker.

All scope notes from the previous approval remain: local actuation interval, not full hold/workflow latency; phase-boundary offsets, not first network-byte timestamps; cold/warm denotes observed power transition, not materialization cache state; procfs representation is milliseconds at centisecond resolution; no claim of general arbitrary-input overflow coverage. This change does not fix or bypass the independent SOL/media/harness holds.

Exact-head witnesses run 36355632694 was still in progress at my latest check. I did not run claim_batch, regenerate YAML, execute the local procfs wet test, or reproduce a full timing-file receipt. Keep normal exact-head and merge-queue gates, and retain actual named-claim execution evidence. No live boot is needed to obtain that evidence.

Source: APPROVE-MERGE at this SHA subject to normal checks. Adjacent measurement limitations: explicit, not global observability closure. Live hardware readiness: existing HOLD unchanged. No dispatch, merge, BMC or hardware action was performed.

@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Sep 27, 2026
Merged via the queue into main with commit c893939 Sep 28, 2026
5 checks passed
@gunbai-bot
gunbai-bot Bot deleted the session/proud-deer-288-relay branch September 28, 2026 04:02
gunbai-bot Bot pushed a commit that referenced this pull request Sep 28, 2026
…the phase clock keeps the SOL supervision before attach and before handoff; HandoffRun carries the collector's last live instant for the watch; regenerate fleet-converge.yml

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
gunbai-bot Bot pushed a commit that referenced this pull request Sep 28, 2026
…rver onto the phase-clocked actuation

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
gunbai-bot Bot pushed a commit that referenced this pull request Sep 28, 2026
…e, pre/post-handoff looks and end look threaded through the timed actuation route; fleet-converge.yml re-derived on main's base (61/111)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant