Repository navigation
Boot Mt. Collins diskless through modeled operations - #10630
Conversation
…false The Mt. Collins unit booted an aarch64 Linux kernel to systemd on 2026-09-05, from an image the controller fetched itself over NFS and presented as a virtual CD, on a machine with NO BOOT DRIVE. Nothing served PXE and nothing ran on the controller. That happened and existed nowhere in the corpus. THE MARKER IS THE BOOTED PROGRAM'S OWN OUTPUT, not the media status. A nonzero redirection_status proves a session was established; serial-over-LAN carried GRUB 2.12, "EFI stub: Booting Linux Kernel...", Linux 6.8.0-71-generic on aarch64, squashfs 4.0, an overlayfs upper, and systemd 255.4 as PID 1. WHAT MAKES IT DISKLESS rather than merely booted: the kernel command line is BOOT_IMAGE=/casper/vmlinuz and casper mounts the squashfs read-only under a tmpfs overlay, so the running root is in RAM. The unit reported 261972592 KiB available, which is why a RAM-resident root is unremarkable here. THE IMAGE DIGEST AGREED WITH THE MODEL'S PIN BEFORE ANYTHING BOOTED IT. The file measured 2ee2163c..., which is exactly extdeps.provisioning.ubuntu_install_media noble_numbat_2404_3_live_server_arm64 .content_sha256 and upstream's own cdimage SHA256SUMS. That is what makes the boot attributable to a known artifact rather than to whatever was on the share. CapabilityOemRemoteMedia IS ADDED, AND THE ANNOTATION BESIDE IT HAD GONE FALSE. It read that the REST remote-media surface "answered 401 -- present, unauthenticated-to, never exercised" and grounded "no capability whatever (nothing was mounted)". True when written; false once an authenticated session attached an image and the host booted it. The row is corrected in place rather than having the capability quietly appended, because the sentence was the reason the capability was absent. The distinction that text drew is still right and is what makes the addition admissible rather than optimistic: a 401 grounds a SURFACE and no capability, because a capability is a claim about a write taking effect. What changed is that the write was performed and its effect observed FROM THE HOST -- the same standard CapabilityBootSourceOverride and CapabilityPowerControl were held to. CapabilityVirtualMedia stays absent on a real distinction: this controller serves no Redfish, so what was exercised is the OEM REST route specifically. CapabilityAccountManagement stays absent because the operation that would exercise it was REMOVED, so nothing can perform the write that would ground it. THREE CONTROLLER BEHAVIOURS ARE TYPED, because a modeled caller that does not expect them will be wrong in the dangerous direction: PUT /api/settings/media/general answered HTTP 500 body code 16425 AND THE SHARE RE-MOUNTED -- the image list, empty before, enumerated both images 23 seconds later. POST stop-media answered HTTP 500 body code 1411 and the start-media that followed succeeded. The share listing is a CACHE, not the share: after a new file appeared the controller returned an empty list including for the file it had previously enumerated, until mount_cd was re-asserted. An operation whose exit arms treat nonzero as refusal will refuse a write that worked, which is worse than a plain failure -- the caller then retries or aborts against a controller whose state has already moved. The success signal for these must be a READ-BACK, the same shape oob_boot_handoff already uses for the boot override. THE CONTROLLER CLOCK IS OFFSET, NOT DRIFTING, and the distinction decides the remedy. Two readings sixteen hours apart give the same offset to the second: 4h03m35s behind at 07:27Z and again at 23:57Z. A clock losing time shows a growing gap; an identical gap means the oscillator is fine and the clock was set wrong, so the remedy is ONE write rather than an NTP dependency. It matters because every SEL entry carries this stamp: the power-cycle receipts this lane relies on are four hours in the past, which does not corrupt them but makes them unjoinable to anything observed from the fleet. No clock-set operation is added -- grounding one needs the write performed and read back, and SurfaceBmcClock remains recorded as a successful READ surface only. megarac.Media IS MODELED AND HAS NO CALLER YET. extdeps.bmc.megarac carried paths and contract notes with no transport, so nothing in this repository could invoke the media surface and the first attach was driven by a hand-written shell script -- the unmodeled realization DESIGN §6 names. The service makes it reachable; the workflow that calls it is not in this commit, so image procurement is NOT yet end-to-end through .dag and this commit does not claim it is. Session auth is cookie-and-CSRF rather than netrc, which is why these cannot reuse the redfish.Http shape beside them, and the credential reaches argv in the session call only -- stated in place rather than hidden, with the next-rung repair named. WHAT THIS BOOT DOES NOT ESTABLISH, recorded on the receipt itself: it is a STOCK live-server image with nothing per-host injected and no identity bound; the backing share is srv2, a FLEET PEER, so the colo case is unproven; and UEFI HTTP Boot remains UNMEASURED because every boot in this lane was forced to a device over IPMI, so the PXE client observation is not evidence that HTTP boot is unavailable. Corpus: 27 blocking, none in any file this commit touches. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D7PLjw5FMYdZ2GWn3FeMpR
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 69f6048592
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| input { bmc_host: NonEmptyStr, cookie_jar: NonEmptyStr, csrf_token: NonEmptyStr, json_body: NonEmptyStr } | ||
| output { body: String from "stdout", success: Bool from "exit_success" } | ||
| transport shell { | ||
| argv: ["curl", "--fail-with-body", "-sS", "-k", "-b", "{cookie_jar}", "-H", "X-CSRFTOKEN: {csrf_token}", "-H", "Content-Type: application/json", "-X", "POST", "-d", "{json_body}", "https://{bmc_host}/api/settings/media/remote/stop-media"] |
There was a problem hiding this comment.
Treat the observed HTTP 500 stop as indeterminate
On the Mt. Collins firmware recorded by this same commit, a successful stop-media returns HTTP 500/body code 1411. Local curl --help all documents --fail-with-body as “Fail on HTTP errors but save the body,” so this command exits nonzero and the following exit map reports failure before a workflow can perform the required read-back. Model this response as indeterminate and determine success from GetRemoteConfigurations, rather than making the known successful response a terminal refusal.
Useful? React with 👍 / 👎.
| input { bmc_host: NonEmptyStr, username: NonEmptyStr, password: NonEmptyStr, cookie_jar: NonEmptyStr } | ||
| output { body: String from "stdout", success: Bool from "exit_success" } | ||
| transport shell { | ||
| argv: ["curl", "--fail-with-body", "-sS", "-k", "-c", "{cookie_jar}", "-X", "POST", "-d", "username={username}", "-d", "password={password}", "https://{bmc_host}/api/session"] |
There was a problem hiding this comment.
URL-encode session credential fields
When either credential contains form delimiters such as &, +, %, or =, passing username={username} and password={password} through -d changes the submitted form fields and causes authentication to fail despite valid credentials. Local curl --help all identifies --data-urlencode as the option for URL-encoded POST data; use it for both named fields instead of raw -d values.
Useful? React with 👍 / 👎.
| operation GetMediaGeneral { | ||
| input { bmc_host: NonEmptyStr, cookie_jar: NonEmptyStr, csrf_token: NonEmptyStr } | ||
| output { body: String from "stdout", success: Bool from "exit_success" } | ||
| readonly | ||
| transport shell { | ||
| argv: ["curl", "--fail-with-body", "-sS", "-k", "-b", "{cookie_jar}", "-H", "X-CSRFTOKEN: {csrf_token}", "https://{bmc_host}/api/settings/media/general"] |
There was a problem hiding this comment.
Add the missing media-configuration write
When the BMC has no mounted share, or when the newly recorded image-cache failure requires mount_cd to be reasserted, this service cannot establish media because its only operation for /api/settings/media/general is a GET. A repo-wide search for that endpoint finds no other executable PUT transport, while the existing contract row says the PUT supplies the server address, source path, share type, and mount_cd; therefore a workflow still cannot reproduce the attach flow this service is intended to expose.
Useful? React with 👍 / 👎.
| // BOOT_IMAGE=/casper/vmlinuz, and casper mounts the squashfs read-only under a tmpfs overlay, so | ||
| // the running root is in RAM. The unit reported 261972592 KiB available, which is why a |
There was a problem hiding this comment.
Do not infer a RAM-resident squashfs from overlayfs
The recorded command line contains no toram option, and the cited boot messages establish only a read-only squashfs lower layer plus an overlayfs upper layer; they do not show that Casper copied the squashfs off the virtual CD into RAM. In the ordinary live-media path the writable upper is memory-backed while the squashfs remains backed by the live medium, so this receipt overstates what the evidence establishes even though the machine is still diskless.
Useful? React with 👍 / 👎.
The previous commit landed service megarac.Media and said plainly that it had no caller, so image procurement was not end-to-end through .dag. This is the caller. EXECUTED AGAINST THE UNIT, from a genuinely detached start rather than a no-op. The image was stopped first (configurations read image_name "" with redirection_status 0), then: ◐ megarac.Media.OpenSession 208ms ◐ megarac.Media.GetRemoteImages 148ms ◐ megarac.Media.StartMedia 104ms ◐ megarac.Media.GetRemoteConfigurations 166ms and an independent read afterwards showed the ISO attached. The shell script that performed the first attach is now redundant. THREE THINGS THE MODELED VERSION DOES THAT THE SCRIPT DID NOT, each one a defect the script had rather than a refinement: THE CSRF TOKEN COMES FROM THE PARSED DOCUMENT. The script matched a regex over the response body — a second reading of JSON beside the one this repository already owns. extdeps.languages.json.parse is the single recognition authority, and a document it cannot parse is now UNREADABLE rather than partially believed: MegaRacSessionTokenUnreadable carries the body that defeated it. THE IMAGE INDEX IS RESOLVED BY NAME, NEVER ASSUMED. start-media takes an image_index and the controller assigns indices by its own enumeration of the share; the same file was index 0 on one scan and index 1 after another image appeared. A baked index attaches whatever now occupies that slot, which on a share holding a memtest ISO and an installer ISO is a SILENT WRONG BOOT rather than a failure. SUCCESS IS A READ-BACK, NEVER AN EXIT STATUS, and this firmware forces it. The previous commit recorded two measured cases of it answering HTTP 500 while the write took effect. An operation treating non-zero as refusal would refuse a write that worked, leaving the caller to retry or abort against a controller whose state had already moved. The confirmation here is the configurations listing naming the image, read separately from the start call. AN EMPTY LISTING IS ITS OWN OUTCOME because it has its own remedy. MegaRacImageListingEmpty means the controller's share cache went stale — measured on this unit, which stopped enumerating even a file it had listed before, until mount_cd was re-asserted. Reading it as "the share has no images" sends an operator to inspect an export that is fine. The listing is a cache, not the share, and the two are different states. MegaRacAttachUnconfirmed is kept distinct from MegaRacStartRefused for the same reason: the mount daemon is asynchronous, so start-accepted-but-not-yet-listed is PENDING rather than failed. It is reported rather than slept through, because a caller that polls until something looks successful has no deadline. THE CREDENTIAL ASYMMETRY IS NAMED RATHER THAN HIDDEN. OpenSession is form-encoded and curl offers no netrc equivalent for a form body, so the password reaches argv for exactly one call; every subsequent operation carries no credential because the cookie is the bearer. The ipmi operations beside this take a password FILE for the same threat and can, because -f exists there. Wrapping the difference away would make one call's exposure look like none. Corpus: 27 blocking, unchanged from the previous head, none in the new files. WHAT IS STILL NOT MODELED: the ISO fetch itself. The image was downloaded to the share by hand-run curl, and nothing in this repository yet places a verified artifact on a controller-reachable share. The digest was checked against extdeps.provisioning.ubuntu_install_media noble_numbat_2404_3_live_server_arm64 .content_sha256 before anything booted it, but that check was also performed by hand. Attach is now modeled; procurement is half. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D7PLjw5FMYdZ2GWn3FeMpR
…rified one
Attach and boot were modeled; procurement was not. The ISO the diskless boot ran
from was downloaded by hand-run curl and checked by hand-run sha256sum, which
left the one step that decides WHAT the machine boots outside the model.
THE ORDER IS THE GUARANTEE, and it is why this is not simply "curl then check".
The fetch lands in a STAGING path outside the export, the digest is measured
there, and only an agreeing artifact is moved into the published path. Fetching
straight into the export and checking afterwards leaves an interval in which a
truncated or substituted image is mountable — and the controller re-scans that
share on its OWN schedule, not this program's, so the interval is not bounded by
anything here. Measurement and publication are reachable only through the arm
where the fetch succeeded, and publication only through the arm where the digest
agreed; both are separate funcs rather than statements after an `if`, the same
shape that keeps the boot handoff's power cycle behind a confirmed read-back.
A MISMATCH LEAVES THE STAGED FILE IN PLACE, deliberately. The bytes that
disagreed are the evidence for the refusal, and deleting them turns a
diagnosable result into a bare assertion. What matters is that they never reach
the published path, not that they cease to exist.
EXECUTED, WITH THE DISCRIMINATING RED FIRST:
wrong pin 3.0 GB fetched, digest measured 2ee2163c..., refused against a
deliberately impossible pin, and the published path was NOT
CREATED — verified from the filesystem, not from the program's own
report. The control publishes to a path the controller does not
read, so a defect in the wall could not have put an unverified
artifact where the unit would boot it.
correct pin fetch succeeded, digest measured, digest AGREED, and publication
failed: /srv/bmc is root-owned and the run is unprivileged. The
existing published ISO is untouched — same size, same timestamp —
so a failed publish did no collateral damage.
THAT SECOND RESULT IS WHY BootImagePublishFailed IS ITS OWN ARM. "The artifact
disagreed with its pin" and "the artifact was fine and I could not write it
where it belongs" have different owners and different remedies; collapsing them
would send someone to re-examine an image that was never the problem. The
refusal names both paths so the remedy is readable from the message.
THE PIN IS CONSUMED, NOT MINTED. extdeps.provisioning.ubuntu_install_media
already carries content_sha256 for this artifact, cited to cdimage's own
SHA256SUMS, and the value here is that figure. When the artifact row grows a
locator field the policy should read both from it rather than carrying either; a
second copy of a digest is a second authority for one fact, and the copy is
always the one that rots.
WHAT REMAINS: the export is root-owned, so the publish step needs an ownership
or privilege decision that is an operational fact rather than a modeling one.
Until it is made, this operation fetches and verifies through the model and
refuses at the last step, which is a better state than the previous one — where
the same work happened with no refusal available at all.
Corpus: 27 blocking, unchanged, none in any file this commit touches.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D7PLjw5FMYdZ2GWn3FeMpR
The modeled fetch verifies an artifact correctly and then cannot publish it: /srv/bmc on srv2 was drwxr-xr-x root root and gunbc runs unprivileged, so the operation ended at BootImagePublishFailed. That refusal was accurate, and the question behind it is not a permissions detail. THE EXPORT IS THE PAYLOAD HALF OF THE BOOT PATH. Write access to the directory a controller reads its virtual media from is equivalent to choosing what the host runs -- the same authority the out-of-band handoff exercises from the control half. This lane spent its effort making the control half admit, verify and refuse; leaving the payload half as "whatever can write the directory" secures one end of a chain. BOTH EASY ANSWERS ARE WRONG IN THE SAME WAY. Escalating the publish step means converge acquires privilege for a file write, which is not proportionate to putting a verified ISO in a directory. Widening the directory to the interactive login is the mirror failure: every process running as that user could then replace the boot payload, and on a workstation-shaped account that population is unbounded and unnamed. THE PRINCIPAL ALREADY EXISTED AND A FIRST DRAFT MINTED A SECOND ONE. That draft carried a free PosixLoginName "gunbc-media"; gunbc.fleet_posix_accounts already owns FleetPosixAccountKey, and FleetAccountAutomation is declared there with login_policy KeyOnlyAutomation -- a non-interactive identity that exists for exactly this. A media-specific account would have been the nicknaming DESIGN §3 forbids. The requirement now names the key. NO DEDICATED BOOT-PAYLOAD ACCOUNT IS SPLIT OFF, though a narrower principal is available in principle. Nothing has established that converge automation should be unable to change what a host boots -- converge is the thing that decides what a host boots -- so the split would be speculative separation. If a fact ever demands it, the key vocabulary is where the variant goes. GROUP AND OTHER ARE BOTH DENIED WRITE, and that is the load-bearing half. An export owned by a named principal but group-writable has MOVED the authority rather than narrowed it: the population able to replace the boot payload becomes the group's membership, which is the unnamed category this row exists to avoid. FOUR STANDINGS, ALL EXERCISED. root is its own arm rather than one case of a wrong owner, because the remedies differ -- a wrong-but-unprivileged owner is a chgrp, while root means the publish step would have to escalate, which is the design decision being refused rather than a misconfiguration to correct. An account outside the fleet roster is its own arm too: reporting it as "wrong principal" would name a required principal while saying nothing about what is actually there. root-owned ExportOwnershipIsRoot the automation principal satisfied an ordinary interactive account wrong-principal an account outside the roster unrecognized The interactive case is the one worth asserting: it is the outcome someone reaches by chgrp-ing the directory to their own login to make the publish step work, which succeeds and silently widens exactly what this row exists to name. THE FIXTURE WAS WRONG AND THE HOST CORRECTED IT. It carried name "gunbc" at uid 997, an account that exists nowhere. srv2's fleet automation principal is gunbc-fleet at uid 1001, already the principal the host reach and identity probes compare against. A fixture whose account does not exist tests the fold against a fiction and passes just as happily; it was found by reading the host rather than by trusting the row. THE HOST CHANGE IS MADE: /srv/bmc is now gunbc-fleet:gunbc-fleet mode 755. No account creation was required. The controller still reads the share and the attached image remains bound (redirection_status 1); the images LISTING went stale again across the change, which is the cache behaviour already recorded in the diskless boot receipt rather than a regression. WHAT THIS DOES NOT CLOSE, and it is a moved gap rather than a solved one: the publish step still refuses from an operator session, because gunbc runs as uid 1000 while the export belongs to 1001. That is the requirement working, not failing -- it says the AUTOMATION principal owns the export, so converge should run as that principal. Special-casing a grant to make an operator session write there would reintroduce precisely the widening refused above. Corpus: 27 blocking, unchanged, none in any file this commit touches. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D7PLjw5FMYdZ2GWn3FeMpR
Three things surfaced while driving the controller for netboot, none of them the object of a measurement, each with an owner outside this lane. They existed only in a session transcript until now. THE UNIT IS RUNNING NON-REDUNDANT, AND IT IS CURRENT RATHER THAN HISTORICAL. PSU2_Status (0xeb) asserts BOTH [Presence detected] AND [Power Supply AC lost] right now — not merely in the event log — so the supply is installed and has no AC input. The corroborating sensors agree: PSU1_POUT 160 W PSU2_POUT 0 W PSU1_FAN_SPEED 5000 RPM PSU2_FAN_SPEED 0 RPM One supply is carrying the entire load while the other is inert. The SEL carries the assertion on 09/04 and again on 09/05 controller time, so this predates the netboot work rather than being caused by it. IT IS RANKED ABOVE THE CLOCK SKEW THIS LANE ALSO RECORDED. A wrong clock makes receipts unjoinable and is fixed by one write. A single-corded machine has no redundancy at all, cannot be fixed by any write, and this lane is actively making that machine something the fleet depends on for boot. The remedy is physical — a cord, an outlet, or a PDU port — and belongs to whoever is in the room. PRESENCE AND AC INPUT ARE SEPARATE FIELDS rather than one health flag, because the present-but-unpowered case is the interesting one: an absent supply is a visible gap in a chassis, while a present one with no cord looks correct from the front and reports itself as installed. Redundancy is a property of the POWERED supplies, not the installed ones. THE CONTROLLER HAS TWO INTERFACES AND ONLY ONE IS THE ONE WE USE. 192.168.1.228 is the dedicated management port; 192.168.1.231 is a second interface on the SAME controller, established rather than assumed — mc info at both addresses reports Device ID 32 and firmware 0.32. It is recorded because it was once MISREAD AS THE HOST. A DHCP capture from that address was initially taken for the host PXE-booting; it was not. Its vendor class is "udhcp 1.21.1", the BusyBox client the controller's own Linux runs, and the host NIC is a third address entirely. The discriminator is the vendor class rather than the MAC's proximity: a PXE client identifies itself as PXEClient:Arch:NNNNN and asks for a boot file, and a Linux DHCP client does neither. A second management address answering the same credentials is also a second reachable surface, which an access policy should know about the unit. THE PROBE EXERCISES BOTH DIRECTIONS. Asserting only that the measured rows are not redundant would pass just as happily against a check that answered "not redundant" unconditionally — and this unit's live case is the unhappy one, so that is exactly the mistake available here. A corrected pair, identical but with AC present, must be reported redundant for the live verdict to carry any information. Corpus: 27 blocking, unchanged, none in any file this commit touches. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D7PLjw5FMYdZ2GWn3FeMpR
briansrls
left a comment
There was a problem hiding this comment.
Verdict: REQUEST CHANGES / HOLD at exact head c1203c9419829360ac928bcb651e14a02ac9b152. Recorded as COMMENT because this connection acts as the PR author and GitHub will not accept a native blocking review. No merge authorization.
First, the operational finding outranks the source review: PSU2 is installed but reports AC lost, 0 W, and 0 RPM while PSU1 carries the machine. Freeze further power-cycle/load work and do not make this host a production boot dependency until someone at the rack restores and verifies the second feed without disturbing PSU1. Ampere's current Mt. Collins GSG calls for both power cables to appropriate power sources.
There is real progress worth preserving: the host-side systemd marker closes “did the aarch64 image execute”; detached media attach is stronger than a no-op; wrong-pin publication refusal is a useful executed discriminator; deferring mksquashfs and reusing FleetAccountAutomation are both the right direction. The following boundaries remain open.
R1 / P1 — The modeled media path reproduces the session leak that just saturated the controller, and its authentication transport is not valid for the password policy. megarac_attach_remote_image opens a session and every outcome returns without DELETE /api/session. The only policy wrapper uses one fixed /tmp/gunbc-mtcollins1-megarac-jar, so the source comment that the jar is attempt-owned is false; concurrent or repeated attempts can overwrite/reuse one bearer file. OpenSession also sends password={password} through raw -d: form delimiters in a valid punctuated password change the submitted field, and the secret appears in argv. -k additionally provides no controller certificate identity. Model one session adapter: attempt-owned owner-only jar, correctly URL-encoded password read from a materialized file, bound controller trust posture, close on every handled exit, verify invalidation, then remove the jar. Session creation or cleanup with an indeterminate transport result needs its own outcome; do not retry a write automatically.
R2 / P1 — Fetch, attach, and boot are three callable surfaces, not one artifact-bound chain. BootImageAdmitted { path, digest } is rendered to bare ProcessExit; mtcollins1_attach_diskless_image independently selects a filename, and the boot handoff independently selects CD-ROM. Nothing makes a successful verified publication a precondition of attaching those bytes, or proves the later boot used that admitted artifact. Carry a sealed attempt/unit/artifact binding from final-path digest readback through exact media configuration and into boot/host acceptance. An independently callable attach of a pre-existing same-named file must not satisfy that chain.
The attach readback itself can fabricate success: it ignores GetRemoteConfigurations.success, substring-searches the raw body for the filename, does not require active redirection_status or the resolved index/session, and returns image_index: 0 regardless of the index it actually selected. A stopped/stale row or an error body containing the name can therefore produce MegaRacAttached. Parse the configurations document, require one exact active binding to the selected image/index, and return that observed index. Malformed/wrong-shaped image JSON is currently collapsed into MegaRacImageListingEmpty; keep unreadable separate from an actual empty cache. A nonzero start response followed by no immediate confirmation is indeterminate/pending on this firmware, not terminal MegaRacStartRefused.
R3 / P1 — The bytes hashed are not structurally the bytes published. Fetch uses one fixed /var/tmp path, hashes it, then calls plain mv to /srv/bmc, with no ownership lock and no digest readback at the destination. Another attempt/process can replace the staging bytes between hash and move. The model also does not establish same-filesystem rename; across filesystems mv may expose a partially copied destination to the controller. Use attempt-owned staging on the same established filesystem but outside the exported namespace, failure-atomic no-replace/authorized replacement, and hash the final path before minting the admitted artifact. The wrong-pin control proves the current branch decision; it does not test this race/publication boundary.
R4 / P1 — Export ownership is declared but neither observed nor enforced. BootImageExportOwnershipRequirement names FleetAccountAutomation and carries owner/group/other write policy, but publish_verified_boot_image never consumes it. export_ownership_standing receives no file mode and never reads those three booleans; its probe is fixture-only. Thus root, an interactive principal, or a group/world-writable export can still publish whenever mv succeeds. Bind an actual stat/effective-principal receipt to the exact export path and require that admission before the first publication write. The source also still says /srv/bmc is root-owned and today's fetch refuses, while the later commit/report says it is now gunbc-fleet:gunbc-fleet mode 755 and publication succeeds; synchronize the live source and evidence rather than leaving both claims active.
R5 / P2 — Four earlier review findings are still open at this head. StopMedia still maps the observed HTTP-500/took-effect response to command failure; there is still no PUT operation for /api/settings/media/general, even though the caller says an empty cache is repaired by reasserting mount_cd; session fields are still not URL-encoded; and the boot receipt/body still infer “root is in RAM” from squashfs + overlayfs without toram or a successful media-detach test. The boot is genuinely diskless and reached systemd. What is not established is independence from the virtual CD/NFS after boot.
R6 / P2 — The newly promoted platform facts have no preserved evidence binding and one predicate overclaims. mtcollins1_platform_observation contains current PSU values and the second-interface identity as authored data but no raw-capture digest/path/attempt. Preserve the actual sensor/SEL and both-endpoint identity responses. power_supplies_redundant proves only that two supplies are present and have AC input; it does not establish health, capacity, or a redundancy state, so rename it to the narrower fact or add the missing inputs. Device ID 32 plus firmware 0.32 is not independently identifying proof that two IPs are the same controller; use a unit-unique FRU/serial or controller-side interface inventory.
Also remove the copied Ubuntu URL/hash authority: this tree already has the artifact row, filename/mirror constructors, and content pin; the Mt. Collins policy should select a mirror and consume that row rather than restating its digest. Update the PR body, which still says Media has no caller and ISO fetch is unmodeled despite the later commits.
Required discriminators include: zero leaked sessions on every result arm; two concurrent attach attempts with distinct jars; punctuation-bearing credential success without secret-in-argv; stale/stopped/wrong-index configuration refusing; final published-path digest bound to the attach; a publication race/cross-filesystem refusal; and zero writes for wrong principal or permissive group/other modes.
Validation boundary: exact-head static source/consumer review plus inspection of the committed/reported execution records. I did not operate the controller or reproduce the boot. Exact-head witnesses run 34009053965 was still in progress when reviewed.
Attaching an already-attached image refused. For a converge step that is the failure mode that matters most, since converge is run repeatedly by construction. The cause was a wrong reading of the controller. /remote/images enumerates what is AVAILABLE to attach, not what the share holds: measured on Mt. Collins with two ISOs exported, the listing named both; after start-media attached one, it returned only the other, and the survivor's image_index moved 1 -> 0. The attached image had moved to /remote/configurations. The earlier reading was that the listing is a cache that goes stale, with the operator-facing remedy "re-assert mount_cd". That remedy appeared to work for the worst possible reason: toggling mount_cd detaches everything, returning the image to the available pool. On a host booted from that medium -- this lane's whole purpose -- the advice removes the running root. It is retracted rather than reworded. So configurations is now read first and the converged case performs no write at all. That is not an optimisation: start-media against an already-presented image is a write whose effect on a booted host is unmodelled, and reading first is how we avoid making it. Also drops the index comparison from the confirmation. It was added reading image_index, which is absent from a configurations row, so every genuine attachment was reported unconfirmed -- caught because the check failed closed. Re-reading it as media_index passed, and was still wrong: media_index is the virtual CD device slot, beside session_index, not a position in the image enumeration. It agreed only because this unit presents one CD device, so both are 0. Joining them is nicknaming, and it would have begun refusing on any controller with a second slot while blaming the image. Verified against the live controller, one input differing: the real image returns success in three calls with no write; an absent name refuses after the two reads, before start. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D7PLjw5FMYdZ2GWn3FeMpR
The fetch staged to /var/tmp and published to /srv/bmc as two independent literals. rename(2) is atomic only within one filesystem; on this host both happen to sit on the same logical volume, so the move was atomic and the module's central claim held. Measured, and by coincidence. Nothing established it. Point either literal at a separate mount -- a tmpfs /var/tmp among them -- and mv silently degrades to copy-then-unlink, publishing a growing partial file into a share the controller re-scans on its own schedule. That is exactly the window the staging split exists to close, reopened by editing a path constant, with no diagnostic anywhere. Staging is now derived from the destination, so the two share a directory by construction, therefore a filesystem, therefore the rename is atomic -- and there is no longer an input whose editing could break it. The suffix is deliberately not .iso: staging now lives inside the export, and the controller enumerates by extension, evidenced by memtest86-iso.zip sitting in this same export unlisted. So a partially-written file is not offerable as media even while being written. Also measures the published bytes again at the final path rather than carrying the staged digest forward as though it answered for them. The staged digest establishes what was fetched; it says nothing about what now occupies the path the controller reads. Reporting one as the other is execution-provenance loss. Verified by execution on a controlled fixture whose content, and so whose digest, was authored rather than measured from the tree: fetch, stage, move, and a second digest at the destination. The new mismatch arm has no executed red and the module says so -- it needs a concurrent writer or a mocked move, so it is a reachable-but-unoccupied guard rather than a decoration. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D7PLjw5FMYdZ2GWn3FeMpR
The fetch module authored a published path ending ubuntu-24.04.3-live-server-arm64.iso. The attach module authored an image name spelled ubuntu-24.04.3-live-server-arm64.iso. Two literals, two files, nothing joining them. They agreed, and nothing required them to. What that admitted is the failure this lane exists to prevent. Edit one and converge verifies the digest of artifact A, publishes A into the export, then asks the controller to present artifact B -- which, if B is in the share, succeeds. Every step reports success, every read-back confirms, and the unit boots bytes nothing checked. There is no arm for it because no single operation is wrong; the defect lives in the seam between two that are each individually correct. The published path is now derived rather than authored, which makes the divergence unrepresentable instead of merely detectable. A check that the path ends with the name would have been the shape to avoid: it concedes the two are separately writable and then validates them, and is satisfiable by editing whichever one the author was not thinking about. Composing the path from the name leaves nothing to compare. The decomposition is the honest one rather than a trick to collapse a pair. A published path genuinely fuses two facts with different owners: where this host's controller reads media from, which moves when the export moves, and which artifact is served, which moves when the release moves. They were only ever one string because someone wrote them as one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D7PLjw5FMYdZ2GWn3FeMpR
boot_image_export_ownership described who may own the boot export and was consulted by nothing. The publish called mv directly, so the program published whenever the operating system's permissions happened to allow it -- as root, or into a world-writable export -- while a probe asserted the policy against fixtures. A requirement with no causal position is a description, and describing a wall is not building one. The reason it had no causal position was that nothing could observe ownership: there was no modeled stat. extdeps.tools.coreutils_stat adds one, folding stat(1) into a typed observation whose unreadable arm is a refusal rather than an ownership of nobody -- collapsing those would make a deleted or unreachable export the most permissive input the admission accepts. parse_int's Absent arm refuses rather than defaulting to 0, because 0 is root: an unparseable uid would otherwise arrive at the admission as the most privileged answer available. The judgment now takes a uid and a name rather than a PosixUser. That is a correction, not a convenience: a PosixUser carries a login shell, which is a passwd fact, while the owner of a directory is a stat fact. Asking for a PosixUser meant the only way to feed this from a real observation was to invent a shell nothing had seen -- fabricating a field to satisfy a type, in the function deciding who may write to the boot export. The check sits after the digest agrees and before the move, so the two refusals stay distinguishable: checking ownership first would make every wrong-pin fetch on an unwritable export report an environment problem and hide the artifact defect. The fetch now takes the export directory and image name separately and composes the path itself, because the admission is about the directory while the write is about the path. Taking both as parameters would let a caller admit one place and write to another -- the same seam defect as verifying one artifact and attaching a different one. Verified live on srv2, four arms of the observation on real paths: /srv/bmc admitted, / refused as root-owned, /home/briansrls refused as the operator principal, a missing path refused as unobservable. And the gate itself: the correct pin against a root-owned export verifies the artifact and is then refused at the write boundary, with the filesystem confirming nothing was published and the staged bytes kept as evidence. The wrong-pin control still refuses on the digest, so the ordering holds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D7PLjw5FMYdZ2GWn3FeMpR
The previous commit retracted an overclaim and, in the same paragraph, made one: it said toggling mount_cd pulls the running root out from under a booted host. That is the mechanism -- the presented medium is withdrawn, and casper reads the read-only lower layer from it on demand rather than from RAM -- but it is not what was observed. The toggle was performed on this unit while it was booted from the medium, and the running installer survived: measured afterwards over serial console, still interactive at its own menu. The controller is configured rmedia_retry_count 3 / interval 15, and the system was idle at a prompt rather than reading, either of which may explain it. So the blast radius is unquantified: not catastrophic, not safe. One observation of one idle host is not evidence that a detach during heavy reads is survivable. The remedy is still wrong for the reason that does not depend on any of this -- it "fixes" the listing by undoing the attachment the caller asked for. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01D7PLjw5FMYdZ2GWn3FeMpR
briansrls
left a comment
There was a problem hiding this comment.
Verdict: REQUEST CHANGES / HOLD at exact head e43eb5508621cc6c4dee5f6d11574b2afb72fc50. This connection acts as the PR author, so GitHub cannot record a native blocking review; this COMMENT carries the request-changes verdict. No merge authorization.
Exact-head run 34028505967 is green, including build, clippy, floor, and the aggregate job. GitHub reports the PR mergeable/clean. That is useful standing, but the changed set adds no test-claim file for these pure judgments, and current main advanced after this run; more importantly, the substantive boundaries below remain open.
Disposition of review 5124032031. There is material progress. Configuration-first lookup makes the converged attach path read-only; attachment confirmation now parses an exact active name/media row instead of substring-searching; the selected image index is no longer fabricated as zero; staging and destination are derived onto one filesystem; the final path is re-hashed; real stat is consulted at the write boundary; the fetch and attach policy share one image-name atom; and the PR body plus .dag receipt now distinguish diskless from RAM-resident. Those portions of the prior review are closed or substantially narrowed. The following are not.
R1 / P1 — Session ownership is still not total, and the new path can still leak exactly the resource that saturated the controller. OpenSession can succeed and return a body whose CSRF token is unreadable; that branch returns SessionReleaseNotAttempted and never calls CloseSession. The session was opened. Conversely, a lost/HTTP-error response from either POST or DELETE does not establish whether the remote allocation/effect occurred. The prose says SessionReleaseUncertain exists, but the type has no such arm; CloseSession.success == false is called Refused even though this firmware is already measured applying writes while returning HTTP 500. The production renderer also drops cleanup failure whenever the attach outcome itself failed (other => other).
The reported live 401 after DELETE is good execution evidence, but the modeled operation does not perform that invalidation readback and never removes the local jar. The attempt is an arbitrary NonEmptyStr, not a safe/exclusive path allocation, and the jar's owner-only mode is unestablished. OpenSession also still sends the password through raw -d password={password}, in argv, under -k; the existing URL-encoding finding remains open and is now more consequential because valid policy-conforming passwords contain punctuation. Build one session-scope adapter: safe attempt identity, create-new owner-only jar, file-backed URL-encoded credential, bound TLS posture, cleanup on every allocated/possibly-allocated path, protected-read invalidation check, and local removal. Preserve release unknown separately from released/refused.
R2 / P1 — The ownership gate observes the directory owner label, not the declared publication authority. coreutils.Stat returns uid/gid and names, but no mode; the admission ignores gid, all three writable_by_* fields, and requirement.export_path. It also never observes the effective principal performing the move. Therefore a root-running process targeting a gunbc-fleet-owned export is admitted; a 0777 export owned by gunbc-fleet is admitted; and an unrelated gunbc-fleet-owned path is admitted despite the requirement naming /srv/bmc. This does not enforce “only FleetAccountAutomation may publish”; it only recognizes one target owner spelling.
Bind the exact non-symlink directory identity, mode/ACL posture, and effective writer to the admission consumed at the first write. Add controls where target ownership is right but actor, path, group-write, or other-write is wrong. Also remove the now-false source paragraph saying nothing consumes this requirement—the later code does consume it, just incompletely.
R3 / P1 — Same-filesystem rename is repaired, but concurrent publication can still put unverified bytes into the live export. Every attempt for one destination uses the same <published>.gunbc-staging path and plain replacing mv. There is no create-new attempt ownership or exclusive publication boundary. A second writer can replace staging after the first hash; the first move can then publish those different bytes. The final-path re-hash will detect the mismatch, but only after the wrong bytes have occupied the controller-visible filename—the safety property was already violated. The staging file is itself inside the exported namespace; “this firmware did not enumerate this suffix in one observation” is not the same as structurally keeping partial bytes out of the export.
Use attempt-owned staging on the same established filesystem but outside the exported namespace, then one authorized/exclusive no-replace or compare-and-swap publication effect. Final-path readback remains valuable, but detection after exposure is not publication admission. The wrong-pin control exercises the ordinary branch; it does not exercise this race.
R4 / P1 — The verified publication still does not causally reach attach or boot. BootImageAdmitted { path, digest } is immediately erased to ProcessExit. Attach consumes only the shared filename, and boot independently selects CD-ROM. Centralizing the name closes the two-literal drift, but it does not make successful publication a prerequisite: a pre-existing or post-verification replacement with that name can be attached and booted without any BootImageAdmitted value ever existing.
Carry a sealed (unit, attempt, upstream artifact identity, final path, final observed digest, export identity) value into attach; carry the exact active media binding into the boot handoff and host acceptance receipt. A shared string is an authority repair, not an effect-chain binding.
R5 / P1 — Attach still turns observed indeterminacy into a terminal refusal and can write over an unreadable configuration state. The source comment correctly says a nonzero start followed by no immediate confirmation is pending, not refused. The implementation still returns MegaRacStartRefused in exactly that case. On this firmware, HTTP 500 can accompany a successful asynchronous write, so that false terminal can provoke a retry against state that already moved.
Also, configurations_confirm_attachment maps malformed or wrong-shaped configuration JSON to false; the caller then consults available images and may issue StartMedia over an unknown presented-media state. Parse and classify the configuration document before any write. A nonzero start with no bounded confirmation needs an indeterminate/pending arm, never Refused solely from curl status. The earlier StopMedia HTTP-500 contract and the missing executable PUT /api/settings/media/general also remain unresolved surfaces introduced by this PR.
R6 / P1 — The evidence receipt no longer identifies the evidence file, and the file carries mutually contradictory live claims. artifacts/bmc/mtcollins1-diskless-boot.json was modified to append the detach observation, but mtcollins1_diskless_boot_capture_digest and mtcollins1_diskless_boot_byte_count remain cefdf6… / 4361. The receipt therefore cannot name the current bytes absent a hash collision. Preserve the original immutable capture and add a separately hashed correction/observation, or recompute the identity and make the provenance change explicit.
The current JSON still says “casper runs the squashfs from RAM” and still records the image list as a stale cache repaired by reasserting mount_cd; later fields in that same file retract both readings. The .dag receipt likewise still carries the stale-cache interpretation while the attach authority says that diagnosis was backwards and its remedy detaches media. A correction appended beside an active contradictory claim does not remove the claim. Separate raw observations from interpretations and leave one current conclusion.
R7 / P2 — One artifact authority is still authored three times. The policy still hard-codes the Ubuntu URL and SHA, and mtcollins1_boot_artifact hard-codes the filename, although extdeps.provisioning.ubuntu_install_media already owns the selected artifact row, filename derivation, mirror URL construction, and content pin. Select the upstream artifact plus mirror and derive all three values; do not retain a local copy while saying the upstream row is authoritative.
The platform rows also still need preserved sensor/SEL and controller-interface evidence. power_supplies_redundant establishes only two present AC inputs, not health/capacity redundancy, and Device ID 32 plus firmware 0.32 is not a unit-unique proof that both addresses are one controller. The accepted one-feed risk can remain an operator decision, but the observations it accepts still need evidence and appropriately narrow names.
Required discriminators before re-review: token-unreadable and ambiguous-open paths cannot disappear as “nothing to release”; cleanup failure is visible beside every attach outcome; released-cookie invalidation and jar deletion; punctuation-bearing file-backed login with no secret in argv; right-owner/wrong-actor and permissive-mode publication refusals; two concurrent publishers cannot expose one another's bytes; attach cannot run without the exact admitted artifact; malformed presented-media JSON causes zero writes; HTTP-500/no-confirmation remains indeterminate; and the committed evidence digest resolves to the actual current bytes. Add durable pure tests for these parsers/folds rather than relying only on reported live invocations.
Validation boundary: exact-head static source/consumer review plus the committed execution records and exact-head CI status. I did not operate the controller or reproduce the wet runs. Current main has advanced since run 34028505967, so required-CI freshness and exact merge composition must be re-established after substantive repair.
| Absent => | ||
| MegaRacAttachResult { | ||
| outcome: MegaRacSessionTokenUnreadable { detail: session.body }, | ||
| session_release: SessionReleaseNotAttempted, |
There was a problem hiding this comment.
P1: OpenSession already succeeded here, so SessionReleaseNotAttempted is not “nothing was opened.” This path never calls CloseSession and recreates the leak when the response is valid enough to allocate a session but its CSRF payload is unreadable. Carry allocated-but-unreleasable/unknown explicitly and invoke the session-scope cleanup/recovery path.
| MegaRacAttached { image_name: image_name, image_index: image_index } | ||
| } else { | ||
| if start.success == false { | ||
| MegaRacStartRefused { detail: start.body } |
There was a problem hiding this comment.
P1: This contradicts the immediately preceding contract. On the measured firmware, a nonzero HTTP result can accompany an applied asynchronous write; no immediate confirmation therefore means pending/indeterminate, not StartRefused. A caller may otherwise retry a write that already took effect.
| Absent => ExportOwnershipUnrecognizedAccount { observed: observed_name } | ||
| Present { value: k } => | ||
| if fleet_account_key_matches(a: k, b: requirement.required_owner) { | ||
| ExportOwnershipSatisfied |
There was a problem hiding this comment.
P1: This admits solely from the target owner's mapped role. requirement.export_path, gid/mode policy, and the effective writer are not consulted. A root process or a process targeting a 0777 gunbc-fleet-owned directory passes even though the stated authority forbids both. Bind exact path + mode/ACL + effective principal at this write boundary.
| // memtest86-iso.zip, whose name contains "iso", and the controller has never listed it. So a | ||
| // partially-written file under this suffix is not offerable as media even while it is being written. | ||
| fn boot_image_staging_path(published_path: NonEmptyStr) -> NonEmptyStr { | ||
| join([published_path as String, ".gunbc-staging"], "") as NonEmptyStr |
There was a problem hiding this comment.
P1: This is one shared staging pathname per destination, not an attempt-owned file. Concurrent calls can replace it after another call hashes it; final rehash detects only after different bytes have been moved into the live exported name. Use create-new attempt-owned staging plus an exclusive/conditional atomic publish.
|
|
||
| fn mtcollins1_render_fetch(outcome: BootImageFetchOutcome) -> ProcessExit { | ||
| match outcome { | ||
| BootImageAdmitted { path: _, digest: _ } => ExitSuccess |
There was a problem hiding this comment.
P1: This erases the only final-path/digest binding. The attach and boot entry points can then succeed using a same-named pre-existing or replaced file without consuming a successful publication. Preserve a sealed BootImageAdmitted-derived value through attach and boot instead of reducing it to ProcessExit.
| "diskless_evidence": { | ||
| "kernel_cmdline": "BOOT_IMAGE=/casper/vmlinuz", | ||
| "memory_available_kib": 261972592, | ||
| "note": "casper runs the squashfs from RAM with a tmpfs overlay; nothing is written to local storage, and this unit has no boot drive to write to", |
There was a problem hiding this comment.
P1: This active field still asserts the squashfs is in RAM, while the appended correction says the lower layer remains media-backed. This file was also modified without updating the receipt's declared SHA-256/4361-byte identity. Keep raw evidence immutable and carry corrections in separately identified evidence, or update one non-contradictory artifact and its identity.
briansrls
left a comment
There was a problem hiding this comment.
Verdict: REQUEST CHANGES / RECUT at exact head e43eb5508621cc6c4dee5f6d11574b2afb72fc50. GitHub records this as COMMENT because this connection acts as the PR author and cannot submit a native blocking review. Exact-head witnesses run 34028505967 is green; this is still not approval or merge authorization.
I agree with the operator's organizational diagnosis. This should not be merged and followed by a cleanup. The branch has produced valuable evidence and several useful mechanisms, but its root decomposition is wrong: reusable MegaRAC realization, generic artifact publication, srv2 staging policy, one-unit observations, operator risk acceptance, and the final Mt. Collins integration are authored together under gunbc.machine_intake. The problem is not the directory spelling by itself; it is that those facts now reach across one another as globals and duplicate authorities that already exist.
Use this branch as quarry and recut before merge.
R1 / P1 — The publisher is ambient-local while claiming to target srv2
gunbc.machine_intake_boot_image_fetch.fetch_verified_boot_image hard-codes LocalExec; mtcollins1_fetch_boot_image carries no staging-host or transport identity. It therefore writes /srv/bmc on whichever machine happens to execute the function, not structurally on srv2. The ownership layer then interprets that local machine's account names through srv2_account_key_for_name.
This bypasses the existing gunbc.machine_intake_staging authority, which explicitly says srv2 is only a candidate realization of an abstract staging service and still records today's realization as StagingUnobserved. A caller's current working host is an undeclared global, not a staging realization.
The reusable publisher must consume an admitted staging realization/host transport and the selected artifact. The srv2 endpoint, export, principal realization and observation belong together in the fleet/staging realization; they must not be imported by the generic operation.
R2 / P1 — A parallel boot-delivery authority was built instead of completing the existing one
gunbc.boot_artifact already owns BootArtifact. gunbc.boot_artifact_delivery already declares itself the one solver and carries the exact target, artifact, staging offer and MegaRacShareOffer. This PR instead adds BootImageFetchOutcome, mtcollins1_boot_artifact, independent fetch/attach entry points, and strings for name/path/digest.
mtcollins1_boot_artifact only centralizes the image name and export directory; it is not the content-addressed artifact. mtcollins1_boot_image_fetch still re-authors the Ubuntu URL and SHA-256 even though extdeps.provisioning.ubuntu_install_media already supplies the artifact, filename/mirror constructors and pin. It then renders BootImageAdmitted { path, digest } down to ProcessExit. The attach path imports only the name, and boot handoff is independent again. A pre-existing or replaced same-named file can therefore be attached without consuming the verified publication result.
Delete the parallel root. Extend the canonical delivery path so a sealed, target/attempt/artifact-bound publication receipt is consumed by MegaRAC attach, whose exact attachment is then consumed by boot handoff and host acceptance. Centralizing two literals is not the same as preserving execution provenance across the chain.
R3 / P1 — The ownership gate does not enforce its own requirement
BootImageExportOwnershipRequirement carries the required path plus owner/group/other write policy, but coreutils.Stat.PathOwnership observes only uid/gid/names, and export_ownership_standing checks only the owner key. It never checks:
- that the observed path equals
requirement.export_path; - owner/group/other mode bits;
- the effective principal performing the publication;
- the host on which the observation was made.
Thus a gunbc-fleet-owned world-writable directory, a different gunbc-fleet-owned path, or a root process publishing into an automation-owned directory can be admitted. The fixture probe cannot close fields the production observation does not carry.
Bind host + exact path/filesystem identity + owner/group/mode + effective publisher into one observation, evaluate an explicit policy supplied to the publisher, and carry the admitted context through the first write. The generic operation must not import one global /srv/bmc requirement.
R4 / P1 — Session cleanup still misses an allocated-session path, and the client contract is not safe for the credential policy
After OpenSession succeeds, a missing/unparseable CSRFToken returns SessionReleaseNotAttempted; no CloseSession is issued. The Mt. Collins renderer then labels that arm “session was never opened,” which is false on exactly this path. This can reproduce the leak that saturated the controller. A lost/ambiguous session-open response is also not equivalent to “nothing allocated.”
SessionReleased is minted from DELETE's exit status alone even though the PR's own acceptance used a subsequent protected read returning 401; that invalidation is not in the modeled operation. The local cookie jar is not deleted, and its path is derived from a free String/NonEmptyStr, not a path-safe attempt identity or create-new owner-only allocation. The negative-control wrapper does not even reject an empty attempt.
OpenSession still sends raw -d password={password} and uses -k. A valid punctuation-bearing rotated password can change form decoding, the password remains in argv, and the TLS peer is not bound to the enrolled controller.
Extract one reusable MegaRAC session realization: path-safe attempt ownership, owner-only/create-new jar, file-backed URL-encoded credential input, explicit controller trust posture, allocation-uncertain and release-unverified outcomes, close on every path on which a session may exist, protected-read invalidation, and local cleanup. The curl cookie-jar realization does not belong in the upstream API authority.
R5 / P1 — An unreadable presented-media document is treated as “not attached,” permitting a write over unknown state
configurations_confirm_attachment collapses JSON parse failure or a non-array document to false. megarac_attach_with_token then reads the available-images list and may call StartMedia. MegaRacConfigurationsUnreadable covers transport failure only, despite the operator-facing text claiming malformed presented state refuses before a write.
Return a typed parse/shape standing and refuse before start whenever current attachment state is unknown. Also, redirection_status != 0 is not an established active-state decoder, and the comment saying “non-zero start followed by no confirmation is pending, not refused” is contradicted by the implementation, which still returns MegaRacStartRefused when start.success == false. The fresh-controller path remains incomplete as well: there is still no modeled PUT for /api/settings/media/general, so this code can reuse an already configured share but cannot establish one.
R6 / P1 — Hash/readback improvements detect a race after publication; they do not prevent exposure
Deriving staging beside the destination and rehashing the final path are improvements over the previous head. The staging path is nevertheless deterministic, shared by every attempt for that destination, and placed inside the exported namespace. Two callers can overwrite the same staging file between fetch, hash and move. A conflicting writer can put wrong bytes at the published path before the final digest detects them, during which the controller may observe them. The claim that .gunbc-staging is unofferable is inferred from one zip not appearing in one listing, not a structural contract.
Use an attempt-owned create-new staging object on the established same filesystem but outside the exported namespace, then an exclusive/expected-current publication boundary and final readback. A lock/CAS/no-replace guard is appropriate here as protection around one legitimate effect writer; it must not be used to excuse two semantic owners. Replacing an artifact already attached to a running host also needs an explicit consumer/current-state decision.
R7 / P1 — The active evidence and the typed interpretation contradict one another
artifacts/bmc/mtcollins1-diskless-boot.json still says Casper runs the squashfs from RAM and still carries the retracted “stale share cache / reassert mount_cd” interpretation. The corrected .dag receipt says diskless, not RAM-resident, yet that same receipt also retains the old cache paragraph. The later appended correction does not make the earlier live fields stop asserting the opposite. Because this artifact is cited by digest as evidence, this is not harmless history; it is one evidence object with mutually incompatible interpretations.
Split immutable raw responses from interpretation receipts. Preserve raw bytes unchanged, author one current interpretation, and let superseded readings live in commit/review history rather than as simultaneously true fields.
The access/capability provenance also needs a clean attempt boundary: mtcollins1_evidence_manifest_digest explicitly covers the earlier read and boot-control write pair, while CapabilityOemRemoteMedia is added from the later diskless-boot artifact; successful_write_surfaces still names only boot control. Create a later observation/receipt or compose the manifests explicitly rather than silently broadening the earlier observation.
R8 / P2 — mtcollins1_platform_observation is four authorities in one file
This module combines PSU sensor observations, controller-interface inventory, an operator risk decision, and fixture checks, with no raw-capture digest/path/attempt. AcceptedPowerRisk is policy, not sensor evidence; the PR contains no evidence reference for the claimed operator decision. If an external operator acceptance exists, bind that receipt explicitly. power_supplies_redundant proves only that two present supplies report AC input; it does not establish health, capacity, independent feeds, or redundancy. The second-interface sameness claim is based on a common device ID/version rather than a unit-unique controller/FRU binding.
Keep the per-attempt boot/access receipt in machine intake. Put physical state into the fleet/placement observation authority, network-controller identity into the access/network observation, and risk acceptance into its decision authority. Do not merely move this mixed module under a different directory.
Recommended recut
- Evidence-only slice: preserve raw captures and author corrected, target/attempt-bound receipts for the diskless boot, MegaRAC behavior and access capability. No generic actuators. Split PSU/interface/risk facts by authority.
- Reusable MegaRAC realization: upstream API shapes stay in
extdeps.bmc.megarac; curl/session/media state machinery lives in a BMC realization module with a complete session lifecycle and fail-closed decoders. No Mt. Collins constants. - Staging/publication realization: extend
gunbc.machine_intake_staging,gunbc.boot_artifact, andgunbc.boot_artifact_delivery; bind a real staging host/transport, explicit publication policy and failure-atomic artifact receipt. No srv2 globals in the generic publisher. - Mt. Collins integration: one small policy/composition slice selects the existing Ubuntu artifact and MegaRAC share offer for the exact unit/attempt, then carries publication → attachment → boot handoff → host marker without dropping to
ProcessExitbetween stages.
The earlier review's final-path readback, ownership-at-boundary, exact attachment parsing, session close, and RAM-residency corrections have all advanced in this head, but only partially; the remaining findings above are why the green run does not authorize placement. Recut at the semantic roots rather than accumulating another repair layer on this branch.
Keep the diskless-boot operations branch current so it can land on today's main.
… treating unreadable presented-media as detached. An opened session with no CSRFToken was reported as never opened, which is the leak that filled this controller. Malformed configurations collapsed to false and could StartMedia over unknown attachment state. The capture artifact also still claimed a RAM-resident root after the receipt had retracted that. Co-authored-by: Cursor <cursoragent@cursor.com>
#10888 landed, so the rows this lane demoted on 2026-09-09 come back with the arms that were missing. No page was re-fetched: the evidence never moved, only the shape that could hold it. databank_ewr2_density CapabilityAtLeast { lower: watt(10000) } + CoolingMediumPublished { AirCooling } digital_fortress_piscataway CapabilityAtLeast { lower: watt(15000) } + CoolingMediumUnspecified njfx_wall_density marketed_real_interval(watt(2000), watt(16000)) + CoolingMediumUnspecified NJFX NOW ASSERTS WHAT ITS PAGE SAYS. The old arm had ONE slot, so the row carried only the 16 kW top and dropped the 2 kW floor the operator publishes in the same sentence - it asserted LESS than its source. Constructed through marketed_real_interval rather than a record literal because the arm's payload is sole_constructor, so the mint is the only route and a reversed pair becomes QuantityUnreadable instead of being silently swapped into order. EWR2 IS NOW THE ONLY SCREENABLE DENSITY IN CENTRAL NEW JERSEY, which the medium coordinate made visible rather than caused. It is the only row whose operator names a cooling medium - the word AIR in "10kW+ Cab Density Air". NJFX and Digital Fortress publish real figures that no medium request can reach. SO THE CONTROLS SPLIT IN TWO, AND THE SPLIT IS THE POINT. w_a_capability_floor_supports_from_below_and_never_bounds_above asserts EWR2 and nothing else, and its third conjunct is the REGRESSION CONTROL: a floor does NOT refuse one watt above itself. The witness this replaces asserted exactly that refusal at 10,001 W against a page promising AT LEAST 10 kW, and was green for two days. w_an_unlabelled_figure_is_inventoried_but_unscreenable asserts TWO facts of NJFX and Digital Fortress - no medium request reaches them, AND they hold one figure each. The second conjunct is what distinguishes them from an unread row: both answer no to every screen, one because nothing was published and one because nothing was said about cooling, and a control that only asked the screen would read them as identical. Evidence, and the compile signal is honestly unavailable: witnesses 47/47 - 6 mine, 41 shared, the shared roster having grown 21 -> 41 with the carrier's own controls, so these rows satisfy the carrier author's witnesses and not only mine. Each re-keyed control was EXECUTED before this commit, which is the rule this lane filed against itself after the carrier author reintroduced a deleted inference while translating a witness. `gunbc compile` cannot go green on this entry: dag/gunbc/machine_intake/megarac_media_attach.dag carries 17 section 4c violations, is byte-identical to origin/main, appears 0 times in this diff, and arrived with #10630. claim_batch scopes to the witness module's real closure, which is why execution evidence exists at all. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016jFtgPtXxTj1kE8wwZUNsG
A second defect from #10630, currently masked: boot_image_export_ownership declares fleet_account_key_label as a total serialization of FleetPosixAccountKey — its own annotation says "exhaustive so that a new account variant is a compile error here rather than a silent non-match" — but matches only seven arms. The type has carried eight since #10714 added FleetAccountNotificationService, and #10630 added this file with the seven it knew. It cannot surface on main today because parse refuses before exhaustiveness is ever asked. It becomes reachable the moment the corpus parses again, so it is fixed here rather than left to red main a second time: this is the designed wall firing, not a new one. The label is mechanical, not a judgement call — every arm is its variant suffix in kebab-case, and the string's only consumer is fleet_account_key_matches, which compares two labels for equality, so it needs to be distinct and conventional and it is both. Every other consumer of FleetPosixAccountKey was checked: fleet_ssh_access, managed_access_bootstrap and posix_principal_allocation_witness_test carry no total match over the key, and fleet_posix_accounts already covers all eight. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NMGuMfzK6MWHJyybTNSxfy
…wn_repair This PR is that class again, six days after the first specimen and with the next-rung trigger still unfired: parse now reports 5391 files parse-clean, and namespace-wave-admission refuses with NotEvaluated because the BASE revision still carries the 17 diagnostics this PR removes. The refusal is correct and the class is unchanged, so this appends a receipt rather than proposing a mechanism. It also records one fact the first specimen could not show: an aborted phase bounds the error total behind it. Parse refused before exhaustiveness was ever asked, so the non-exhaustive fleet_account_key_label match #10630 also introduced was invisible for the entire outage and became reachable only once the parse repair was applied. The class therefore conceals its own scope, which is why a repair author cannot report main fixed on one phase going quiet. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NMGuMfzK6MWHJyybTNSxfy
…red since #10630) (#10932) * Hoist megarac_media_attach annotations to module-item grain #10630 landed dag/gunbc/machine_intake/megarac_media_attach.dag with three annotation blocks sitting inside function bodies. Only module-item grain is modeled (DESIGN §4c), so the parse phase refused with 17 errors, and the namespace-wave-admission phase then reported "no head index" because no index could be built over a corpus that does not parse. Main has been red on every head since. The three blocks carry real rationale, so they move above the declaration each describes rather than being deleted: the allocated-session leak onto megarac_attach_remote_image, the unreadable-presented-state arm onto megarac_attach_with_token, and the StartRefused reservation onto megarac_start_and_confirm — folded into that function's existing paragraph, which already stated the pending-not-refused half. No other .dag file in dag/, src/v1 or src/v2 carries a body-position annotation. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NMGuMfzK6MWHJyybTNSxfy * Complete fleet_account_key_label; the eighth account variant had no arm A second defect from #10630, currently masked: boot_image_export_ownership declares fleet_account_key_label as a total serialization of FleetPosixAccountKey — its own annotation says "exhaustive so that a new account variant is a compile error here rather than a silent non-match" — but matches only seven arms. The type has carried eight since #10714 added FleetAccountNotificationService, and #10630 added this file with the seven it knew. It cannot surface on main today because parse refuses before exhaustiveness is ever asked. It becomes reachable the moment the corpus parses again, so it is fixed here rather than left to red main a second time: this is the designed wall firing, not a new one. The label is mechanical, not a judgement call — every arm is its variant suffix in kebab-case, and the string's only consumer is fleet_account_key_matches, which compares two labels for equality, so it needs to be distinct and conventional and it is both. Every other consumer of FleetPosixAccountKey was checked: fleet_ssh_access, managed_access_bootstrap and posix_principal_allocation_witness_test carry no total match over the key, and fleet_posix_accounts already covers all eight. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NMGuMfzK6MWHJyybTNSxfy * Record the second measured specimen of fail_closed_gate_refuses_its_own_repair This PR is that class again, six days after the first specimen and with the next-rung trigger still unfired: parse now reports 5391 files parse-clean, and namespace-wave-admission refuses with NotEvaluated because the BASE revision still carries the 17 diagnostics this PR removes. The refusal is correct and the class is unchanged, so this appends a receipt rather than proposing a mechanism. It also records one fact the first specimen could not show: an aborted phase bounds the error total behind it. Parse refused before exhaustiveness was ever asked, so the non-exhaustive fleet_account_key_label match #10630 also introduced was invisible for the entire outage and became reachable only once the parse repair was applied. The class therefore conceals its own scope, which is why a repair author cannot report main fixed on one phase going quiet. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NMGuMfzK6MWHJyybTNSxfy --------- Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…in_offers_single_cabinet (#10968) main is RED on the declarations check: required-ci: declarations FAIL IMPORT-MEMBER-ABSENT dag/test/claim/colo_nj_central_south_census_test.dag:13:51 `test.claim.colo_nj_central_south_census` imports `grain_admits_single_cabinet` from `test.claim.colo_cabinet_density_witness`, which declares no such name NEITHER CONTRIBUTING CHANGE IS WRONG IN ISOLATION, and the composition is the defect. #10865 (a675b70, merged 19:45:49Z) renamed `grain_admits_single_cabinet` to `grain_offers_single_cabinet` in the shared witness module and updated every call site it could see. #10864 (de7622e, merged 19:46:19Z) imports the old name, which existed when that PR was written. The two merged THIRTY SECONDS APART, both MERGEABLE/CLEAN with all four required checks green, and neither floor verdict could contain the other's change. This is the third instance today of one class: a stranded caller, produced by two changes whose verdicts were each concluded, correct, and computed against a base that excluded the other. The earlier two were the megarac §4c break (#10630) and the `mutation_status_is_commit_ambiguous` rehoming (#10925 wrote the edge, #10923 deleted its target). A 30-second gap is the sharpest form: waiting longer for a floor to conclude does not help when both floors HAD concluded. The repair follows the rename rather than reverting it. `grain_offers_single_cabinet` carries the identical signature `(g: RetailGrainStanding) -> Bool` and the identical body — it folds `retail_grain_single_cabinet_wording` to a Bool — so this is a pure spelling change at eight call sites, and the new name is the one the shared module now declares. EVIDENCE - GREEN by execution: `gunbc compile --entry dag/test/claim/colo_nj_central_south_census_test.dag` -> rc=0, `0 blocking error(s)`, `compiled: 61 files emitted`, zero IMPORT-MEMBER-ABSENT. The completion marker is quoted beside the count deliberately: a count with no completion marker beside it is a claim about the pipeline rather than about the subject. - RED control, on the real acceptance path: main's own required floor on de7622e, run 34522373789, reporting this exact finding. The failing content is the parent of this commit. Both contributing lanes are archived, so this was cut by a blocked lane rather than routed. Claude-Session: https://claude.ai/code/session_01998xs4ojJKWGptxcuxVWHN Co-authored-by: Brian Searls <briansearls1@gmail.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Mt. Collins boots diskless, and the steps that get it there are modeled rather than driven by hand.
The unit has no boot drive. It now boots Ubuntu from an ISO served over NFS from srv2 and presented as virtual media by its MegaRAC controller, and every step of that — fetch the image, verify it, publish it into the export, attach it, hand off the boot device — is a
.dagoperation with a typed refusal, not a shell script someone ran once.What each piece does
Fetch (
boot_image_fetch,mtcollins1_boot_image_fetch) downloads to a staging path, measures the digest there, and moves into the export only on agreement with the pinextdeps.provisioning.ubuntu_install_mediaalready carries. The staging path is derived from the destination, so the two share a directory and therefore a filesystem, which is what makes the rename atomic. It was previously two independent literals that happened to land on one mount — true on this host, and unenforced. The published bytes are then measured again at the final path rather than carrying the staged digest forward, because those answer different questions.Attach (
megarac_media_attach,extdeps.bmc.megarac) opens a controller session, resolves the image index by name rather than assuming one, starts the media, confirms against the configurations listing, and closes the session on every path including refusals. Success is never an exit status here: this firmware answers HTTP 500 on writes that take effect, so the read-back decides.Ownership (
boot_image_export_ownership,extdeps.tools.coreutils_stat) decides whether the publish may happen at all, at the write boundary, from a realstat.Boot handoff and the platform/boot receipts record what was actually observed on the unit.
Three defects this found in its own work
A dangerous remedy, retracted.
/remote/imagesenumerates what is available to attach, not what the share holds: attach an image and it leaves that listing for/remote/configurations, and the survivors reindex. I had read this as a stale cache and written the operator-facing remedy "re-assertmount_cd". That remedy appears to work for the worst possible reason — togglingmount_cddetaches everything, returning the image to the available pool. On a host booted from that medium, the advice removes the running root. It also made the operation non-idempotent, which for a converge step is the failure that matters most. Configurations is now read first and the converged case performs no write at all.A nicknaming join in the confirmation. The index check first read
image_index, which is absent from a configurations row, so every genuine attachment reported unconfirmed — caught only because the check failed closed. Re-reading it asmedia_indexpassed, and was still wrong:media_indexis the virtual CD device slot, besidesession_index, not a position in the image enumeration. It agreed only because this unit presents one CD device, so both are0. The comparison is gone.Two authorities for one artifact. The fetch module authored a published path ending
ubuntu-24.04.3-live-server-arm64.iso; the attach module authored an image name spelling the same thing. Two literals, two files, nothing joining them. Edit one and converge verifies artifact A, publishes A, then asks the controller to present B — every step succeeding, every read-back confirming, and the unit booting bytes nothing checked. There is no arm for it because no single operation is wrong.mtcollins1_boot_artifactnow owns the name and the export directory as two atoms, and the path is composed from them.The ownership requirement now decides something
It previously described who may own the boot export and was consulted by nothing — the publish called
mvdirectly, so the program published whenever the OS permissions happened to allow it, while a probe asserted the policy against fixtures.The reason it had no causal position is that nothing could observe ownership: there was no modeled
stat. There is one now, and its unreadable arm is a refusal rather than an ownership of nobody — collapsing those would make a deleted export the most permissive input the admission accepts.parse_int'sAbsentarm refuses rather than defaulting to0, because 0 is root.The judgment also stopped taking a
PosixUser. That carries a login shell, which is a passwd fact, while a directory's owner is a stat fact — so the only way to feed it from a real observation was to invent a shell nothing had seen, in the function deciding who may write to the boot export.Verification
Executed against the live controller and the live host, not asserted.
401on the released cookie./srv/bmcadmitted,/refused as root-owned,/home/briansrlsrefused as the operator principal, a missing path refused as unobservable.Not claimed
BootImagePublishedDigestMismatchhas no executed red. Producing one needs a concurrent writer or a mocked move, so it is a reachable-but-unoccupied guard rather than a decoration, and the module says so with its next-rung trigger.toramis absent from the observed command line, so the read-only lower layer is read from the presented media on demand. An earlier revision of the receipt claimed the root was in RAM; it does not follow from what was observed, and the difference decides whether the host depends on the boot server for as long as it runs.admin/admindefault stands; IPMISet User Passwordis refused by this firmware and the REST path returned "attempt to set previous password" against a policy-conforming secret. The spare account is contained (Enabled User IDs: 1).🤖 Generated with Claude Code
https://claude.ai/code/session_01D7PLjw5FMYdZ2GWn3FeMpR