Skip to content

A microVM guest registers only under its own attempt label, and the shakedown job it serves exists - #12119

Closed
gunbai-bot[bot] wants to merge 52 commits into
mainfrom
session/quiet-koi-746-attempt-labels
Closed

gunbai-bot[bot] wants to merge 52 commits into
mainfrom
session/quiet-koi-746-attempt-labels

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

Stacked on #12011. Covers step (1) of the real-run plan: the microVM guest can no longer register under the host's production labels, and the job it exists to serve now exists.

What was wrong

The JIT mint derived its labels from gunbc.runner_registration_labels, which produces [self-hosted, linux, arm64, gunbc-heavy, <host>]. That is the label set fleet-desired.yml and the srv1-pinned fleet-converge jobs target, so the first real attempt could have taken a production converge job inside the shakedown VM. The existing positive control pinned exactly that set.

What changes

  • gunbc.runner_jit_mint: the labels are derived from the attempt alone. jit_config_request_for no longer takes a label list. A registration carries self-hosted plus its own runner name, which is already the attempt's unit name. A caller cannot hand in the host set. The LabelsUnresolved refusal arms can no longer be reached, so they are deleted.
  • gunbc.runner_microvm_shakedown_workflow emits .github/workflows/microvm-shakedown.yml (a new generated artifact). It is a workflow_dispatch job with runs-on: [self-hosted, ${{ inputs.attempt }}]. Its known task hashes a fixed payload with the guest's sha256sum. The job fails unless the result equals the digest extdeps.crypto.sha2 derives from the same payload in the model (ccb509dd…, cross-checked against a real sha256sum). It writes one result line to the job summary: attempt, runner, revision, run. There is no checkout and no action. The attempt input reaches the script through env, never by interpolation.

Not closed here, declared

GitHub's runner agent adds default labels (self-hosted, OS, architecture). Whether a JIT registration also gets the OS and architecture labels is not modeled. If it does, the guest would also match [self-hosted, linux, arm64], which witnesses.yml, heal.yml and one fleet-converge job use. That arm is closed by the runner group, not the label: microvm_jit_runner_group_restriction_frontier names a group restricted to this workflow file, which is also why the job is its own file. That frontier blocks every real attempt until it lands. How the group gets created is escalated.

Evidence (local, fresh claim_batch at this head)

  • 9/9 JIT witnesses in github_app_registry_witness_test PASS. The new witness_jit_mint_carries_no_host_production_label refuses any host, class, OS or arch label; its red is the previous set.
  • 2/2 in runner_microvm_shakedown_workflow_witness_test PASS. Its injection claim's red is the first writing of the script.
  • generated_artifact_gate main_wet regenerates the tree with no drift beyond the new file and its .gitattributes row.

🤖 Generated with Claude Code

Brian Searls and others added 30 commits September 21, 2026 20:54
…nPID, the 28 GiB shakedown slot row, the reservation-authorized planner arm, and the controller's release locus

Deliverable 1 of the microVM slot end-to-end item: gunbc-microvm-slot@.service
(Type=exec, root, ExitType=main, KillMode=control-group) execs
gunbc.runner_microvm_slot_controller microvm_slot_controller_main, which composes
the readiness store, the host CAS reservation re-read, the cell binding, the
JIT mint, the per-slot guest shape, the reservation-authorized launch plan and
the durable network receipt into one call of run_controller. The
controller_main_pid_consumer_frontier row is dissolved per its trigger.

The unit and the controller's release locus (gunbc + sources + Firecracker +
jailer, root-owned) are installed through the live_deploy release locus,
parameterized by owner rather than mirrored, as a new fleet-converge mode
microvm_controller_install (a new HostEffect arm, an apply wet entry with
readback).

Deliverable 2: the srv1 shakedown slot (srv1-13) carries a 28 GiB MemoryMax,
MemoryHigh at the ceiling and MemorySwapMax=0 through
gunbc_runner_slot_desired_for, consumed per slot by the cell digest, the slice
directives, the sudoers grants and the guest shape; the fleet row does not
move and the expecting-red production-shape witness stays red. The width cost
is derived in memory_admitted_width_of_envelope and the change is declared as
gunbc.rung_drop microvm_shakedown_slot_row_above_fleet_row.

Parent decision 2026-09-21: the planner gains a sealed, trust-domain-selected
reservation-authorized arm (no production route constructs an ExecutionGrant),
with the grant producer declared as its frontier. Two further frontiers are
declared where owned: the App control-plane observation the mint needs, and
the admin-edge ConvergedSlotNetwork receipt at
runner_microvm_converged_receipt_path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…-install job

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nted mover (run 35655377521 refused rsync as the job user)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…stall names its source relative to the job cwd

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ROOT), and both release-locus units set it from one row

Parent ruling 2026-09-21: cli_run's workspace_root_from scaffold named its own
dissolution -- release bins receive checkout-root at spawn -- and the first start of
gunbc-microvm-slot@srv1-13 on srv1 (exit 101, 'not inside a git checkout') is the
population that needs it; the approval broker's serve unit is dead on srv1 for the same
reason. One env name, read at one site (workspace_root_resolve, consumed by
workspace_root and process_workspace_root); when set it is the authority and the walk
is not consulted, refusing typed unless the path holds dag/ and the locus's tree
receipt; when unset the walk is exactly as before. Two executed controls in
workspace_root_discovery_tests. v1 admission (gunbc.v1_maintenance_standing): this
serves the emitted-binary door the seed already scaffolded, not v1 growth.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…controller; one Filesystem authority; dissolve markers on the two piped-install strings

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… cell (the seed refuses an unlimited cgroup: HostBudgetUnreadable on srv1)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…s installed and the write is checked, the stop slack is a declared row

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ed at 8 GiB on srv1

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ced slice (the seed filled each bound and was killed at it on srv1)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ric_execution_slot_identities, not a count

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…w plus cache headroom (24 GiB); the planning-budget request is dropped as inert (RSS filled 8 and 12 GiB bounds alike on srv1)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…es a body annotation)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…lice (parent ruling 2026-09-22); the controller figure is homed in the width authority and cites the attribution lane's frontier rather than being tuned to the symptom

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…t discharged

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…an answer for UnmeasuredRoot

Review 69931 (DESIGN §5, fabricated default): whole_corpus_compile_public_root_demand_peak
matched the demand row and answered byte_size(0) on the UnmeasuredRoot arm. That arm is
unreachable only while the roster row stays MeasuredForRoot, and this change made the figure
load-bearing twice over -- it sizes the controller slice's MemoryMax and it is a term in
gunbc_microvm_shakedown_slot_charge_bytes, which host_memory_admitted_width derives srv1's
admitted slot count from. A zero there is not a stop: 0 + 8 GiB is a plausible cgroup bound
(measured fatal -- OOM-killed at 7.9G under an 8 GiB bound) and a 16 GiB drop in the charge
hands srv1 two extra admitted slots.

There is no non-fabricated answer for that arm, so the fix is to have no arm. The measured
figure is hoisted to whole_corpus_compile_public_root_peak, the demand row is BUILT from it,
and the one process-sizing consumer reads the row directly. If this root ever becomes
UnmeasuredRoot the row is deleted with the arm and every consumer refuses to compile rather
than reading a default.

Receipt: gunbc run --source-root dag --source-root src/v2 --entry
dag/test/claim/runner/runner_microvm_slot_controller_witness_test.dag --function
the_controller_unit_joins_its_own_bounded_no_swap_slice_outside_the_cell -- the closure
typechecks and the claim evaluates true (the host's trailing refusal is the ProcessExit
wrapper convention, not the claim).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ion in the roster

The floor refused this PR's changed-witness set with an inhabitance error at two call sites in
test.claim.typed_argv_exec_realization_witness_test: `privileged_commands.skip(n: 4).first()`
produces Optional<PrivilegedCommand> and was handed to a `cmd: PrivilegedCommand` parameter.

The defect is latent on main, not written here; what this change did was reach it. The roster is
derived (grounded_principal_privileged_commands maps the principal's sudo_grants), so altering
fleet membership and the sudoers projection pulls this witness into the floor's changed-witness
set. Checked before attributing it to the wall: gunbc#12046's floor is SUCCESS on the same wall
and gunbc#12045 refuses it in that PR's own changed file, so the wall is sound and this site is
simply now within reach.

Index 4 was the deeper fault and the type error was its symptom. A positional citation into a
derived list re-points silently whenever the roster moves, and it never said WHICH grant was
meant -- both claims assert the answer is the tailscale one. The selector now asks for it by
command_path against fleet_tailscale_binary_path, the same authority the roster derives the path
from, and each claim matches Present/Absent with Absent returning false: a roster that no longer
carries the grant fails the claim rather than skipping it.

Receipt: gunbc run --source-root dag --source-root src/v2 --entry
dag/test/claim/typed_argv_exec_realization_witness_test.dag --function
witness_tailscale_roster_selects_grant_list_not_execute_probe -- the closure resolves with no
inhabitance error and the claim evaluates true, so the selector finds the grant rather than
passing through the Absent arm.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…he seed, so the seed stops transcribing them

Review 69990 (DESIGN §3/§5): the spawn-root env name and the tree-receipt name each had two
authors -- the .dag rows the emitter renders into the release-locus unit, and a `const` of the
same spelling hand-written in cli_run.rs -- and the only thing standing over them was
`gunbc_workspace_root_env_name == "GUNBC_WORKSPACE_ROOT"`, a row compared to a copy of its own
value. That check stays green when the row is edited, which is precisely §5's tell. If either
side moved alone every release-locus unit would exit 101 in workspace_root at start, the failure
the spawn arm exists to repair.

NOT the prescribed remedy, and the substitution is deliberate. The review asked for a witness
reading cli_run.rs, on the precedent of the hand-Rust equivalence witnesses. Two reasons against:
a check over two independently-authored literals is a drift DETECTOR that fires after the fork
already exists, and a new tree-reading witness is a new WET witness, which
v2.workflow.floor_changed_witness blocks for any identity the PR touches (RouteGapBeforeVerdict
and DeclinedDiscoveryExcluded both map to PlannedWithoutTerminalVerdict) -- so the prescribed
form likely cannot reach a verdict on this PR at all.

So the fork is removed instead of watched. gunbc.release_locus_seed_constants_emit projects both
rows into release_locus_seed_constants_generated.rs, registered in the generated-artifact roster
and its drift wall, and cli_run.rs consumes them rather than declaring its own. The join is now
the Rust compiler plus the drift gate: a .dag spelling moves the seed at the next regen, and a
hand edit of the generated file is refused. Construction over validation, and the same argument
gunbc.evaluation_budget_consequence_emit already records for refusing a runtime read of the .dag.

The surviving claim asserts the projection CARRIES each row's value. The emitter never spells
either literal, so it goes red if a row moves and the projection does not.

Receipts, both executed:
- ctrl-build --remote -- bash -lc 'ls -la src/v1/stage0/src/release_locus_seed_constants_generated.rs
  && cargo check -p v1-compiler --lib' -- the file is present (500 bytes) and the seed compiles
  clean against the generated consts. The ls is a control: an earlier run reported Finished over
  an UNTRACKED generated file that the mirror could not have carried, and that green proved
  nothing.
- gunbc run --entry dag/test/claim/cli_run_workspace_root_hand_rust_witness_test.dag --function
  workspace_root_spawn_arm_names_its_env_receipt_and_controls -- true.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…arms the new artifact variant owed

TWO THINGS, the second a repair of the first commit in this pair.

1. workflow_dispatch_choice_input_projects_options_list discriminates at ONE interface --
workflow_trigger_entry turning a WorkflowDispatch with a choice input into a yaml mapping whose
`mode` carries an `options` sequence. It reached that interface by forcing
fleet_converge_workflow and taking its first trigger, which constructs every DispatchInput's
description (one is a multi-thousand-character string) for a fact about none of them: 528572 eval
steps against a 72300 budget, rising with every input any lane adds. It now supplies the trigger,
which is the construction two sibling claims in the same file already use.

The predicate was ALSO fusing two subjects, which is why supplying alone did not work: it
hard-wired fleet_converge_mode_options, so it asserted the PROJECTION and the CONTENT at once.
Only the first is a fact about workflow_trigger_entry, and the second is true by construction --
the workflow literal writes `options: fleet_converge_mode_options`, so comparing against that row
restates what the source already carries (DESIGN §5), and the roster's own coherence is
established by every_dispatch_option_is_a_wire_value_of_the_vocabulary. The expected list is now
a parameter, so the discrimination is intact: all options reach the yaml as strings, none dropped
and none added.

The pairing obligation is named on the carrier: the real producer is still run end to end by
fleet_converge_workflow_has_build_job_needs_release_bins and the rendered-yaml claims, so
deleting them would leave this supplied input unpaired.

2. ReleaseLocusSeedConstantsGeneratedRsArtifact left two GeneratedArtifact matches non-exhaustive
-- gunbc.ledger_row_coherence and gunbc.instruments.docs_projection_gate -- so the previous commit
does not resolve. The exhaustive-match wall is what caught it; nothing I ran before pushing
reached either module, which is the defect in the receipt rather than in the wall. Both arms are
added.

Receipt: gunbc run --source-root dag --source-root src/v2 --entry
dag/test/claim/workflow_dispatch_input_witness_test.dag --function
workflow_dispatch_choice_input_projects_options_list -- the closure resolves and the claim
evaluates true. The discriminating red is recorded in the same pair of runs: with the predicate
still hard-wired to the fleet roster the supplied two-option trigger returned FALSE, so the
parameter is load-bearing rather than decorative.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…authorable red, and the third non-exhaustive match

Review 70028, both findings, plus the resolution break that was the actual CI blocker.

1. THE CENSUS ROW WAS FALSE. cli_run_workspace_root_discovery_loc_delta_net binds
`v1_compiler.cli_run workspace_root_from` at WholeDeclaration and stood at 12 -- the
.git-ancestor kernel alone -- while this change grows exactly that declaration. A row left at 12
is the parallel ledger DESIGN §6 refuses: it states a figure the tree contradicts, and its only
consumer (`... > 0`) cannot see it move. Re-derived to 70, and the INSTRUMENT is named on the
carrier rather than the count alone (§6): added non-blank lines of this scaffold's hand-Rust in
cli_run.rs outside its #[test] module, 35 + 20 + 3 = 58, with the 42 test lines excluded on the
same basis the original 12 excluded its own. 12 + 58 = 70.

2. TWO CONJUNCTS WERE DECORATION INSIDE AN OTHERWISE REAL CLAIM, and the review is right that
this is the same objection the claim's own note raises one conjunct earlier.
`string_contains(s: <row>, pattern: "walk_is_not_consulted")` compares a row to a slice of its own
value and stays green after the seed test it names is renamed or deleted. Deleting them outright
would have left both discriminator rows with no consumer at all (§3c), so the claim now asserts
the scaffold names FOUR DISTINCT, non-empty controls. That has an authorable red and it was RUN:
setting the refusal row to the authority row's value -- the copy-paste that makes an unwritten
control look covered -- returns false; restored, it returns true.

What the replacement does NOT establish is stated on the carrier: that the seed declares those
names. No claim here can, because a witness reading cli_run.rs is a new WET witness and
v2.workflow.floor_changed_witness blocks any such identity its own PR touches.

Also corrected: this change's own de-fork had made the scaffold's prose false -- it still said
"the seed transcribes both spellings" after the seed had stopped transcribing them.

3. THE FLOOR BLOCKER WAS A THIRD NON-EXHAUSTIVE MATCH, generated_artifact_emit
artifact_extra_valid, a SECOND match over GeneratedArtifact in a file I had already edited once.
My earlier search grepped a sibling variant while excluding the files I had touched, which is why
it found two sites and not three. The five sites are now enumerated by arm count;
generated_workflow_provenance carries `_ => none` and needs nothing. The cost rows this run was
meant to re-measure were never reached, so those figures are still owed.

Receipt: gunbc run --source-root dag --source-root src/v2 --entry
dag/test/claim/cli_run_workspace_root_hand_rust_witness_test.dag --function
workspace_root_spawn_arm_names_its_env_receipt_and_controls -- true, and false under the planted
duplicate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…fold that cannot says why

Review 70050. The sharpest form of the finding is the one worth answering: this change MINTS a
ByteSize authority (whole_corpus_compile_public_root_peak) and then unwraps it on the next line to
add a bare Int, which gunbc.runner_microvm_slot_unit re-wraps one module later. std.measure
measure_add is the existing surface for adding two measures of one quantity and scale, so reaching
past it is DESIGN §2's re-invention of a concept that already has a home. That the sibling
gunbc_runner_slot_memory_max_bytes sits on main in the same file is prevalence, not precedent.

SUMMED IN THE CARRIER, where the carrier is this change's to move:
- microvm_controller_slice_memory_max() returns ByteSize via measure_add over the peak row and
  microvm_controller_slice_cache_headroom (now a ByteSize row). The pass-through in
  runner_microvm_slot_unit that re-wrapped the Int is DELETED rather than retyped -- two names for
  one fact is the §3 fork, and the module imports the authority directly now.
- gunbc_microvm_shakedown_slot_memory_max is a ByteSize row; the two slice directives that wrapped
  it with byte_size(...) read it directly.
- gunbc_microvm_shakedown_slot_charge() returns ByteSize via measure_add.
- microvm_slot_unit_timeout_stop() returns Second via measure_add over the five bound rows. Second
  is Measure<Time, One, Nat>, so the same surface applies verbatim; the projection to a scalar moved
  to microvm_slot_unit_timeout_stop_sec, which is the one place the rendered directive needs it.

TAGGED RATHER THAN CONVERTED, once, with the reason on the carrier: memory_admitted_width_of_envelope
divides by gunbc_runner_slot_memory_max_bytes, an Int row standing on main and not this change's to
move. Converting the fold alone would take a ByteSize, unwrap it to divide by that row, and re-wrap
-- the same carrier drop relocated rather than removed. It carries
`🟡 dissolve-on: feature:fleet-slot-row-on-a-measure-carrier`, naming measure_fit_count_floor as the
surface that owns the quotient and its zero-divisor guard once the fleet row carries ByteSize. This
is the second arm the review itself offers.

Two annotations that cited the renamed symbols are corrected in the same change; a citation that no
longer resolves is a §3 defect, not a cosmetic.

Receipt: all four claims over the changed arithmetic execute true --
the_controller_unit_joins_its_own_bounded_no_swap_slice_outside_the_cell,
the_unit_stop_timeout_covers_the_controllers_declared_bounds,
the_shakedown_slot_carries_28_gib_no_swap_and_no_throttle_line_while_the_fleet_row_stands,
the_shakedown_slot_is_charged_its_cell_plus_its_controller_before_the_fleet_row_divides. Running all
four rather than only the one naming a renamed symbol is deliberate: twice in this lane a receipt
whose closure did not reach the consumers missed a break that the floor then found.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ledger-Repair-Judged: docs/design-rung-drops.md
Heal-Candidate-Run: 35699456339
…with the new generated artifact

THE LAST COST BLOCKER. fleet_converge_workflow_has_build_job_needs_release_bins is about the job
partition and its needs edges, and it read them off fleet_converge_workflow -- which forces the
WHOLE row, every job's steps AND every DispatchInput description under `on`, one of which runs to
thousands of characters, to establish seven ids: 521845 eval steps against a 72300 budget,
identical across two separate runs (so the claim's own deterministic work, not contention), and
rising with every job or input any lane adds.

The seven jobs were written inline in the Workflow literal, which made the workflow the only route
to the fact. They are now `fleet_converge_jobs()` in the owning module and the workflow reads it.
That is a single authority with the workflow as one consumer, not a test affordance bolted on: the
roster is where the jobs are declared, and there is no second list to drift.

THE COMPANION FIX IS SECOND-ORDER AND IS WHY THE GATE WAS RUN AT ALL. Registering
ReleaseLocusSeedConstantsGeneratedRsArtifact and emitting its .rs was not sufficient: .gitattributes
is ITSELF derived from the artifact roster, so a new artifact moves it too, and the committed file
was stale by exactly the one line that binds the new path to the refusing merge driver. Regenerated
through main_wet on the gate rather than hand-edited, which is what the finding itself instructs.

Receipts:
- gunbc run --entry dag/test/claim/workflow_dispatch_input_witness_test.dag --function
  fleet_converge_workflow_has_build_job_needs_release_bins -- true.
- gunbc run --entry dag/gunbc/instruments/generated_artifact_gate.dag --function main -- the only
  finding was .gitattributes. fleet-converge.yml is ABSENT from it, which is the receipt that
  matters for the roster move: lifting the jobs into a function is emission-neutral, so the
  committed workflow yaml is unchanged.
- The eval-step figure for the roster-reading form comes from CI; this change does not claim it
  clears 72300 until that run reports it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…he one its shape suggests

Review 70080. Three annotations -- gunbc.live_deploy.emit release_locus_tree_copy_steps,
gunbc.runner_microvm_slot_unit, gunbc.runner_microvm_slot_controller -- justified the root-owned
locus as buying "a tree root execs that the job user cannot write" / "root must not exec a binary
the job user can write". On srv1 that protection does not hold, and the tree already says so:
ghrunner holds NOPASSWD over a PATH-UNRESTRICTED /usr/bin/install on every host carrying
/etc/sudoers.d/gunbc-deploy, install(1) is the very mover the install step elevates with, and the
class is rostered at gunbc.rung_drop the_job_user_holds_a_path_unrestricted_root_grant and
gunbc.recurring_failure_mode the_grant_a_frontier_refuses_is_already_held_from_another_authority.

That is DESIGN §4b(1) inflation in this change's own prose: a class's rung is the MINIMUM across
its in-scope paths, and these sentences cited the strongest while the granted-mover path stayed
silent. It was also inconsistent with this same diff twelve lines away, where the guest-image note
states the shakedown makes no isolation claim rather than implying one.

Each carrier now separates the two facts. WHAT THE LOCUS BUYS: the unprivileged-by-default write
path is removed -- nothing lands bytes there without a granted privileged mover, so an ordinary job
step, a stray tool or a mistaken relative path cannot -- and the unit execs this tree's VMM rather
than the job user's HOME copy. WHAT IT DOES NOT BUY: a locus the job user cannot write, because the
same principal reaches it through the same grant. The standing drop is named at each site, with the
statement that this change adds no grant and removes none, so the stronger claim becomes true when
that drop retires rather than when this locus lands.

Comments only; no declaration, value or emitted byte moves.

Receipt: gunbc run --source-root dag --source-root src/v2 --entry
dag/test/claim/runner/runner_microvm_slot_controller_witness_test.dag --function
the_slot_unit_execs_the_controller_entry_as_main_pid_with_the_declared_teardown -- true, and the
annotation-grain sweep over the three files reports no body-grain block.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…racted against the measurement

MEASURED, AND THE HYPOTHESIS IS FALSIFIED. The job roster was extracted to bring
test.claim.workflow_dispatch_input_witness fleet_converge_workflow_has_build_job_needs_release_bins
under its floor budget, on the reading that `on` -- whose DispatchInput descriptions run to
thousands of characters each -- dominated the row. Two required-floor runs:

  whole Workflow row : 521845 eval steps
  roster             : 521510 eval steps
  budget             :  72300

335 steps, 0.06%. The cost is constructing the seven Job values themselves, step bodies included;
a List<Job> projection forces every one, so reading the roster cannot avoid it. The suspect was
wrong.

The extraction is KEPT, on the ground that survives: the jobs were inline in the Workflow literal,
which made the workflow the only route to the job partition, and the roster is now that authority
with the workflow as one consumer. What is retracted is the cost claim -- both annotations now
carry the two figures and state the refutation, because an annotation asserting a saving this lane
measured away is the same overclaim review 70080 flagged in this change's prose one commit ago.

What would actually be cheap is named rather than implied: a roster of ids and needs edges ALONE,
consumed by the seven job functions for those two fields. That restructures all seven and is not
attempted here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…, consumed by the seven builders

Parent ruling, item 2. Each of the seven job builders hardcoded its own `id` and `needs`, so the
graph -- which jobs exist, and which must wait for the release binaries -- was authored in seven
places and owned nowhere. FleetConvergeJobEdge is that graph, and all seven builders now take their
needs from fleet_converge_job_needs(id:); no literal needs list survives in a builder.

THE LOOKUP REFUSES RATHER THAN INVENTS. An absent id would otherwise fall to an empty needs list,
which is the one wrong answer that still emits: a host-touching job with no edge to `build`
installs whatever binaries happened to be on the host. FleetConvergeJobEdgeAbsent carries the id
and the needs projection renders a REFUSED line instead of a dependency.

WHY THIS IS NOT THE PREVIOUS ATTEMPT REPEATED. Lifting the JOBS into their own list bought 335
eval steps of 521845, because forcing a List<Job> forces every Job including every step body --
the projection WAS the construction, now filed as gunbc.recurring_failure_mode
extracting_a_collection_whose_elements_are_the_cost. These edges are ids and lists of strings built
independently of any Job, so a consumer reading the graph forces no job body. That is a difference
in mechanism, stated before the measurement rather than hoped for after it.

The claim reads the edges; its inhabitance obligation stays with the rendered-yaml claims in the
same file, which run the real workflow end to end -- so what is established is the graph the
workflow EMITS, not the graph a builder was told.

Receipts:
- gunbc run --entry dag/test/claim/workflow_dispatch_input_witness_test.dag --function
  fleet_converge_workflow_has_build_job_needs_release_bins -- true.
- RED, run: dropping the microvm-controller-install edge row fails the claim. Reported honestly --
  it fails as an evaluation error on the now-absent index rather than a clean false, so the count
  conjunct is carrying part of that red.
- gunbc run --entry dag/gunbc/instruments/generated_artifact_gate.dag --function main -- ZERO
  findings, so fleet-converge.yml is byte-identical with every needs list now computed through the
  lookup. That is the receipt that matters: the refactor is emission-neutral.
- The eval-step figure comes from CI. This change does not claim it clears 72300 until that run
  reports it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review 70098. outcome_is_success was an exact re-mint of gunbc.command_runner
process_outcome_admitted -- same three arms, same verdicts -- in a module that already imports from
that very module, so the canonical name was one import member away. DESIGN §3: a second name for
one concept, and §2's test that net concepts must not grow by re-invention. The canonical function
even documents this exact caller shape (a `test -f` whose product is the exit status and whose
stdout nobody wants), so the re-mint did not even carry a different rationale.

Deleted, with both call sites -- path_is_executable and path_exists -- now asking
process_outcome_admitted.

The extdeps.shell.exec import went with it: ProcessOutcome and its three variants were reachable
from this module ONLY through the deleted function, so keeping the line would trade a §3 fork for
an unused import. (The neighbouring bare `import extdeps.shell` is pre-existing and untouched.)

Receipt: gunbc run --source-root dag --source-root src/v2 --entry
dag/test/claim/runner/runner_microvm_slot_controller_witness_test.dag --function
the_controller_selects_only_a_fabric_member_out_of_its_instance_name -- the closure resolves with
the import removed and the claim evaluates true.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… dying

Parent ruling: a control that DIES is weaker evidence than one that ANSWERS false, because an
evaluation error is also what a typo, a renamed field or a broken fixture produces -- so the red
stops discriminating the thing it was built for.

THE PRESCRIBED REMEDY DOES NOT APPLY HERE, and the evidence was already in hand: the count
conjunct was ALREADY first (`edges.length() == 7 && edges[0].id == ...`) when the control died on
the absent index. `&&` in this evaluator does not spare its later operands, so no reordering turns
that into a false.

So the claim is TOTAL instead -- it reads no index at all. The projected id list is compared
against the expected spelling (count, membership and order in one structural comparison) and the
edge property is folded with `all`: build needs [], every other row needs [build]. Same content as
the indexed form, nothing weakened.

Receipts, both run:
- green: the claim returns true.
- red: dropping the microvm-controller-install edge row returns FALSE -- a clean answer, where the
  indexed form produced an evaluation error under the same edit. An eval error from this claim now
  means something else genuinely broke, which is the separation the ruling asked for.
Roster restored to seven rows after the control.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@gunbai-bot
gunbai-bot Bot marked this pull request as ready for review September 23, 2026 05:49
…edown script carries its bash-emit marker (review 70414)

plan_jit_mint and dispatch_jit_mint take a MicrovmRunnerGroupStanding instead of a
bare id; RunnerGroupRestrictionUnobserved refuses (JitMintRefusedRunnerGroupUnrestricted
-> DispatchRefusedRunnerGroupUnrestricted). The slot controller passes the unobserved
arm, so the block microvm_jit_runner_group_restriction_frontier states is now the
match itself. New red witness: an otherwise-authorized mint refuses on it.
microvm_shakedown_known_task_script gains its dissolution row naming bash-emit (#5828),
rendered as the script's first line like its ci_spec siblings.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 23, 2026

Copy link
Copy Markdown
Contributor Author

Both findings in review 70414 held; fixed in 7340d67.

  1. Bash-emit marker. microvm_shakedown_known_task_shell_emit_dissolution_trigger names bash-emit (General orchestration intent to Bash emit fold over grammar rows #5828) and is rendered as the first line of the emitted script, like the ci_spec siblings.
  2. The frontier's block was prose, not code. You were right: plan_jit_mint authorized a mint into any group. It now takes a MicrovmRunnerGroupStanding (RunnerGroupRestrictedToShakedownWorkflow | RunnerGroupRestrictionUnobserved) instead of a bare id. The unobserved arm refuses (JitMintRefusedRunnerGroupUnrestricted, carried through dispatch_jit_mint as DispatchRefusedRunnerGroupUnrestricted). runner_microvm_slot_controller mint_attempt_credential passes the unobserved arm today, so every real attempt refuses at the mint until an observer produces the restricted arm. The frontier row now names that match as its consumer. New red: witness_jit_mint_refuses_a_runner_group_not_observed_restricted (owned cell, sufficient credential, required binding; only the group can refuse). The positive controls pass the restricted arm.

Local: 10/10 JIT witnesses and 2/2 shakedown-workflow witnesses PASS on a fresh claim_batch; the slot controller typechecks; main_wet regenerated the yml.

— sent from quiet-koi-746

…bash_emit, not concatenated shell (review 70427)

The concat script and its bash-emit dissolution row are deleted: the capability the
row named exists, so the scaffold was never needed. The step is built from
bash_build constructors and rendered by bash_emit_stmts; a serializer refusal
becomes a generation refusal. The comparison puts the report in the THEN arm and
the refusal in ELSE, so an empty digest (bash_emit has set -e, not pipefail)
refuses. The result line reads only the runner's env, so no expression reaches the
script. Executed locally: held arm exit 0 with the summary written; with sha256sum
absent, the refusing arm, exit 1.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 23, 2026

Copy link
Copy Markdown
Contributor Author

Review 70427: fixed in 740137f by finishing the construction rather than asking for scaffold approval.

The capability my dissolution row named already exists: v2.extdeps.languages.bash_build constructors folded by v2.workflow.bash_emit. So the concat script and its row are both deleted. microvm_shakedown_known_task_statements builds the step as a bash node tree, and bash_emit_stmts renders it. A serializer refusal becomes MicrovmShakedownWorkflowGenerationRefused, so no yml is emitted from a refused program.

Two consequences of using the real emitter:

  • It supplies set -e but not pipefail. The comparison is therefore [ observed = expected ] with the report in the THEN arm and the refusal in ELSE. An empty or erroring digest takes the refusing arm; the inverse spelling would let an erroring [ skip the refusal.
  • The result line reads only the runner's own env (GITHUB_SHA, GITHUB_RUN_ID, GITHUB_RUN_ATTEMPT, RUNNER_NAME, and the attempt via step env), so the rendered script carries no ${{ }} at all. The witness now asserts that over the rendered text.

Executed, not just rendered: the emitted script run locally takes the held arm (exit 0, result line written to GITHUB_STEP_SUMMARY). With sha256sum absent from PATH it takes the refusing arm (exit 1, observed= empty). 2/2 shakedown-workflow witnesses PASS.

— sent from quiet-koi-746

…on: §4c admits module-item grain only (reported by bright-dove-353)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Base automatically changed from session/quiet-koi-746 to main September 23, 2026 19:14
…'s own changes (0491988..0573097)

The stack base landed as squash 0095b48, so the pre-squash history conflicted in
every file #12011 touched. Resolution: main's tree, then this branch's own diff
re-applied three-way; the staged diff against main is exactly those 13 files.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… step, not the rendered script's bytes (floor: 247,268 eval steps over the 72,300 budget)

The claim rendered the whole script through v2.workflow.bash_emit to search it,
paying the emitter on every run for a result shape it only inspected (DESIGN 3, a
witness discriminates at one interface). It now supplies the script empty and
checks that the step binds MICROVM_SHAKEDOWN_ATTEMPT to the attempt input: 52 eval
steps. The real render is executed by the generated-artifact gate on the committed
microvm-shakedown.yml.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
gunbai-bot Bot pushed a commit that referenced this pull request Sep 23, 2026
…n_wet_one, sha256 3fbae87f)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Sep 23, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Sep 24, 2026
@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Sep 24, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Sep 24, 2026
@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Sep 24, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Sep 24, 2026
Brian Searls and others added 2 commits September 24, 2026 11:41
…s YamlEmitRefused arm (main's #11730 replaced serialize_yaml)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
gunbai-bot Bot pushed a commit that referenced this pull request Sep 24, 2026
…one, sha256 cbe53868)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… per-file claim (no refusals, zero uses)

main's census requires every executed workflow file to carry a roster row and a
per-file claim; the new shakedown workflow had neither, so the floor's roster join
failed. Its claim asserts zero uses rather than census_holds's uses > 0: the job is
action-free by design, an unreadable file refuses rather than reading as empty, and
a uses: added later reds it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Sep 24, 2026
@gunbai-bot

gunbai-bot Bot commented Sep 24, 2026

Copy link
Copy Markdown
Contributor Author

Closing as landed via #12178 (59aaecb). #12178 had this branch merged into it, and every file this PR changes is now identical on main: the attempt-only JIT labels, the runner-group standing, the shakedown workflow and its census claim. Nothing is left to merge here.

— sent from quiet-koi-746

@gunbai-bot gunbai-bot Bot closed this Sep 24, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to a manual request Sep 24, 2026
gunbai-bot Bot pushed a commit that referenced this pull request Sep 24, 2026
…diff, applied three-way

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
briansrls pushed a commit that referenced this pull request Sep 24, 2026
…(first approval-capability effect; stacked on #12155) (#12172)

* WIP PR-B: approval-routed microVM runner group ensure

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* PR-B: annotations at module-item grain; escape the path braces in a refusal string

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* PR-B: one annotation block above the service

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* PR-B: regenerate fleet-converge.yml for the microvm_runner_group_ensure mode (main_wet_one, sha256 5c3e0d9f)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* PR-B: regenerate fleet-converge.yml over the rebuilt #12119 base (main_wet_one, sha256 3fbae87f)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* PR-B: the submission-MAC fetch refuses by name, not by the chmod that follows it (review 70675)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* PR-B: the named MAC-fetch refusal is this mode's prelude only; the keyring converge's shared-text line is not this PR's (regen: one hunk)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* PR-B: regenerate fleet-converge.yml over #12119 at e803d33 (main_wet_one, sha256 cbe53868)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Regenerate fleet-converge.yml over main + the stack (main_wet_one, sha256 6807a921)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* microVM controller reads the org's runner groups under its installation token and mints only into the observed restricted group (stacked on #12172) (#12179)

* PR-C: the controller reads the org's runner groups under its installation token and passes the sealed observation to the mint

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* One runner-groups read for every consumer: delete the CLI list and the JSON decode, route inspection, ensure and controller through read_runner_groups (review 70883)

The gh-CLI ListRunnerGroupsJson fetched GitHub's default page and never compared total_count,
so the ensure could read the shakedown group as absent on a page boundary and plan a duplicate
create while the controller's REST read refused. Now there is one operation (REST, per_page 100),
one fold over its outcome (runner_groups_of_rest_read) and one projection with the completeness
and element refusals. perform_organization_jit_mint matches only on jit_mint_generate_gate, so
its pre-token arms are decided once and it carries no unreachable repeat.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Regenerate fleet-converge.yml over main + the stack (main_wet_one, sha256 6807a921)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: gunbai-bot[bot] <289086189+gunbai-bot[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants