Skip to content

Isolate the Rust toolchain homes in required CI, so concurrent runner slots stop rewriting each other's binaries - #9275

Merged
briansrls merged 11 commits into
mainfrom
toolchain-home-isolation
Aug 26, 2026
Merged

briansrls merged 11 commits into
mainfrom
toolchain-home-isolation

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

Line stop. The required evidence path was corrupting itself, so this lands ahead of the mixed-membership OOMD work.

The outage

host symptom
srv4-11 spawn rustfmt: Text file busy (os error 26) — build lane dead
srv3-11 same, 7 minutes later, different PR
srv1-15 cargo: command not found (127) while rustc --version resolved fine

The compiler's own refusal named the cause: the tool "was resolved at admission from PATH … --version was executed successfully before this phase began. So it has been removed, replaced or made unusable while this run was executing; this is not a missing or broken installation."

Each host runs many runner slots under one /home/ghrunner, and every witness job began by installing rustup into it. One slot's install rewrites a toolchain another slot is executing.

This applies a proven in-repo pattern, it does not design a fix

fleet-converge has wiped and set HOME, CARGO_HOME, RUSTUP_HOME under RUNNER_TEMP since someone hit this before — its pin step is literally named "isolated RUSTUP_HOME has no default toolchain". witnesses.yml had no such step and called setup-rust-toolchain into the shared default. Its own diagnostic gave it away: ls -l "${CARGO_HOME:-$HOME/.cargo}/bin", where the :- fallback is only reachable because CARGO_HOME is unset.

Why a third module rather than importing the fleet step

ci_isolate_toolchain_step carried three concepts:

concept commands wanted here?
toolchain-home isolation rm -rf, HOME, CARGO_HOME, RUSTUP_HOME yes
fleet build tuning MAKEFLAGS, CARGO_BUILD_JOBS no
cache policy sccache binding no

CARGO_BUILD_JOBS derives from the selected runner target's RAM-speed budget, so importing it wholesale would impose a number computed for a different subject — a build-policy change smuggled inside an outage fix.

So the intent is decomposed, and gunbc.toolchain_workflow_steps holds only what both workflows genuinely share. The witness workflow does not import the fleet prelude (borrowing a cross-job artifact architecture it doesn't have) and does not duplicate the shell strings (§3 nicknaming).

The composed fleet pipeline is byte-identical by construction — same seven steps, same order, same builder. That was the discriminating control, and it's verified by execution: regenerating produced zero toolchain/isolate/rustup/CARGO_HOME/MAKEFLAGS/sccache line changes in fleet-converge.yml.

The capability closure caught a real fork

I did not predict this one. Adding the pin step made the_live_witness_floor_job_closes_its_capabilities FAIL: the pin consumes RustupCapability, and the witness toolchain step declared only [Cargo, Rustfmt] while fleet declares [Cargo, Rustc, Rustfmt, Rustup] for the identical setup_rust_action. One action, two capability stories.

Corrected the under-declaration rather than weakening the consumer — the action really does install rustup and rustc. It went unnoticed only because nothing consumed Rustup here until now.

The diagnostic now fails closed

The ${CARGO_HOME:-$HOME/.cargo} fallback made a missing isolation step look like a legitimate shared-home mode. It now refuses with ToolchainHomesNotIsolated when either home is unset, so the condition this PR fixes cannot silently return.

Separately found, deliberately excluded

CORRECTION (2026-08-26). At authoring time this branch independently observed generated workflow drift in .github/workflows/fleet-converge.yml. #9287 has since regenerated and merged that projection, so the claim that it is stale on main is no longer true — re-measured against origin/main 7de77a4, the regen actuator produces zero .github/workflows drift. This PR deliberately contains no fleet-converge secret-hygiene change, and no separate projection cut is owed.

Test plan

YAML regenerated from the emitter, not hand-edited — I rebuilt the toolchain binary (3m) rather than hand-apply, because adding steps to two jobs is not the deterministic substitution a text swap is.

claim_batch --hermetic --entry dag/test/claim/workflow_capability_closure_witness_test.dag — 4/4 PASS:

  • the_live_witness_floor_job_closes_its_capabilities (the subject)
  • the_live_fleet_converge_build_job_closes_its_capabilities (the control — proves the refactor preserved fleet)
  • a_job_consuming_cargo_with_no_provider_refuses · a_provider_ordered_after_its_consumer_does_not_close (the ordering REDs)

Emitted order, both jobs: Checkout → Isolate toolchain homes → Install Rust toolchain → Pin rustup default → Build the witness fold.

Not included: #9268's spawn-observation producer. This cut makes cross-job mutation impossible; #9268 later observes per-spawn byte coherence.

What landed after the body above was written

fleet-desired.yml carried the same defect and it is closed here (fdd77154e4). Its job installed a toolchain into srv1's shared /home/ghrunner for as long as it has existed, while the remedy sat one import away. It now emits Isolate toolchain homes -> Install Rust toolchain -> Pin rustup default.

The obligation is now derived, not remembered (gunbc.toolchain_home_standing). A remedy that merely exists is reachable by whoever recalls it, which is exactly how fleet-desired went unfixed. The module folds Workflow.jobs and classifies each job PrivateMutableHome or ImmutableToolchain off its emitted steps — matched against the isolation authority's own command rows and the cited setup_rust_action, never a text pattern — and refuses on a shared-home install, an isolation with no install, or an undecidable expression runner. It is conjoined into all three emitters, so a job added to any generated workflow without a standing produces no yaml at all. The denominator is the emitted Workflow value rather than a roster of job names, so a new job enters the check by construction.

Ten witnesses, 10/10 PASS, including the four discriminating REDs (shared-home install, isolation after install, isolation without install, expression runner) and one live-consumer witness per generated workflow.

Review 56168 was addressed structurally rather than by receipt: the Bool admission predicate is deleted, not annotated — admit_workflow_toolchain_homes returns a typed verdict carrying the failing job_id and its ToolchainHomeRefusal, so each cause renders its own remedy instead of three emitters naming one static string. The two shell carriers now hold separate dissolution obligations, because they wait on different terminal effects: the isolation shell waits on a typed filesystem-and-environment effect, the pin shell on a toolchain-selection effect, and the first could land leaving the second standing.

Scope: this closes the toolchain race, and nothing else

This PR does not fix the floor-lane cancellations, and no reviewer should infer that it does.

Private CARGO_HOME/RUSTUP_HOME touch the path namespace. They stop concurrent runner slots from rewriting each other's toolchain binaries — spawn rustfmt: Text file busy, cargo: command not found while rustc --version still resolves. That is a real outage and it is what this closes.

The floor-lane cancellations are a separate and currently UNDIAGNOSED class. What is established is only this: a required lane can be cancelled mid-build with no verdict about the tree — observed on this PR's own run 32942735743, floor killed at 1m19s during Build the witness fold with Terminate orphan process: cargo / rustc at cleanup, and the sibling build/witnesses jobs reporting 2s and 3s failures downstream of that cancel rather than independently.

Why the cause is stated as unknown rather than as memory pressure. An earlier revision of this section attributed the class to shared host RAM, on a report of two srv3 instances cancelled at ~2460–2520s with heavy swap activity. That attribution is withdrawn — since it was made, srv3 completed a floor run successfully (#9309, srv3-03, 57m25s), which weakens the host-memory story rather than supporting it, and no instrument in this PR ever measured it. A mechanism quoted without an instrument that could refute it is a story, and this PR should not carry one.

So: the toolchain race is diagnosed and fixed here. The cancellation class is real, orthogonal to the path namespace this change touches, and not yet explained by anyone.

Predicting the consequence rather than leaving it to be discovered: once this lands and the ETXTBSY noise stops, cancellations will look newly mysterious — the same lane will keep dying with the loudest concurrent explanation removed. That is the expected outcome of correctly fixing one of two independent causes, not evidence this one regressed.

… slots stop rewriting each other's binaries

LINE STOP. The required evidence path was corrupting itself. rustfmt refused with
ETXTBSY on srv4-11 and srv3-11 seven minutes apart, killing the build lane on two
unrelated PRs, and srv1-15 resolved rustc while cargo had vanished (exit 127). The
compiler's own refusal named the cause exactly: the tool "was resolved at admission
from PATH ... and ran there: --version was executed successfully before this phase
began. So it has been removed, replaced or made unusable while this run was executing;
this is not a missing or broken installation."

CAUSE: each host runs many runner slots under one /home/ghrunner, and every witness job
began by installing rustup into it. One slot's install rewrites a toolchain another slot
is executing.

THE CONSTRUCTION ALREADY EXISTED. fleet-converge has wiped and set HOME, CARGO_HOME and
RUSTUP_HOME under RUNNER_TEMP since someone hit this before -- its pin step is literally
named "isolated RUSTUP_HOME has no default toolchain". witnesses.yml had no such step and
called setup-rust-toolchain into the shared default. So this is applying a proven in-repo
pattern, not designing a fix.

WHY A THIRD MODULE. ci_isolate_toolchain_step carried THREE concepts: home isolation,
build tuning (MAKEFLAGS, CARGO_BUILD_JOBS) and cache policy (sccache). Importing it
wholesale would have imposed a CARGO_BUILD_JOBS derived from the SELECTED runner target's
RAM-speed budget on a job with a different subject -- a build-policy change smuggled
inside an outage fix. So the intent is decomposed into ci_toolchain_home_isolation_steps
and ci_fleet_build_tuning_steps, and gunbc.toolchain_workflow_steps holds the two steps
both workflows genuinely share. The witness workflow does not import the fleet prelude
(it would borrow a cross-job artifact architecture it does not have) and does not
duplicate the shell strings (that is the §3 nicknaming violation).

THE COMPOSED FLEET PIPELINE IS BYTE-IDENTICAL BY CONSTRUCTION -- same seven steps, same
order, same builder -- and that was the discriminating control on the refactor. Verified
by execution: regenerating produced ZERO toolchain, isolate, rustup, CARGO_HOME,
MAKEFLAGS or sccache line changes in fleet-converge.yml.

THE CAPABILITY CLOSURE CAUGHT A REAL FORK, which is the part I did not predict. Adding
the pin step made the_live_witness_floor_job_closes_its_capabilities FAIL: the pin
consumes RustupCapability and the witness toolchain step declared only
[Cargo, Rustfmt], while fleet declares [Cargo, Rustc, Rustfmt, Rustup] for the IDENTICAL
setup_rust_action. One action, two capability stories. The under-declaration is corrected
rather than the consumer weakened -- the action really does install rustup and rustc --
and it went unnoticed only because nothing consumed Rustup here until now.

THE DIAGNOSTIC NOW FAILS CLOSED. It printed ls -l "${CARGO_HOME:-$HOME/.cargo}/bin",
whose :- fallback made a MISSING isolation step look like a legitimate shared-home mode.
It now refuses with ToolchainHomesNotIsolated when either home is unset, so the condition
this PR fixes cannot silently return.

YAML REGENERATED FROM THE EMITTER, not hand-edited: I rebuilt the toolchain binary (3m)
rather than hand-apply, because this change adds steps to two jobs and hand-application
at that structure is not the deterministic substitution the budget-tree.md case was.

SEPARATELY FOUND, DELIBERATELY NOT INCLUDED: .github/workflows/fleet-converge.yml is
STALE ON MAIN against its own emitter -- an unrelated secret-hygiene change (auth headers
moved from -H arguments into files) is authored but never regenerated. My heal run picked
it up; I reverted it so this PR does not absorb somebody else's change under a toolchain
title. It needs its own cut.

Witnesses: workflow_capability_closure_witness_test 4/4 PASS including the live witness
floor job, the live fleet-converge build job (the control), and the two ordering REDs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor Author

(Rewriting this comment — my first version was posted through a shell string where backticks ran as command substitution, so five backticked identifiers were silently eaten and several sentences lost the names they were about. Content below is what I meant to post; nothing in the substance changed.)

Good catch on the mechanism, and I checked it rather than reasoned about it. review 56043's read of the roster is right; the conclusion changes because the gate is inert.

consumers of ci_cost_floor_exempt_step_names
  (excluding its own note and the roster registration):   NONE
callers of group_verdict outside materialization_ladder.dag:   NONE

Nothing computes a cost-floor verdict over workflow step names any more. The roster's own note still describes the machinery it fed — ci.yml, the ci job, the emit_determinism job, group_verdict ordering — and every one of those was deleted in the 2026-08-15 floor cut. So plurality cannot be tripped by this change, because plurality is not computed.

Which is why I am not adding "Isolate toolchain homes" to the roster. An entry there is a claim that the script is below the cost floor, checked by a gate. Adding one to a gate that no longer runs would manufacture an acceptance nothing verifies — it would read as coverage in exactly the way DESIGN §4b calls worse than absent, and the next reader would find a rostered exemption and conclude the step had been assessed. It has not been; there is nothing left to assess it.

The real finding underneath your comment is that the roster is orphaned. ci_cost_floor_exempt_step_names survives as four hand-written strings, still enrolled in gunbc.roster_registry, guarding nothing. It passes the cited-symbol census because the declaration exists — that census asks whether the name resolves, not whether anything reads it. This is precisely the class roster_registry_visibility_note documents at length from three separate incidents: a registration outliving the roster it registers, going quiet rather than red. I am not fixing it here — it is unrelated to a toolchain outage and belongs with the rest of the floor-cut residue — but it should not stay invisible, so it is written down.

One thing your comment makes worth stating explicitly. My module comment says I chose a distinct step name (Isolate toolchain homes, not Isolate toolchain dirs) to avoid silently enrolling a differently-scoped script under an existing acceptance. That reasoning was about collision; you are pointing at plurality, which is the other direction, and which I had not considered. Both land in the same place only because the gate is dead. If that gate is ever revived, this roster needs the new name added and the pin step's plurality re-assessed — the pin name is already rostered but now runs at two further sites. Recording it so a revival does not silently inherit a gap.

— sent from warm-tern-755

Brian Searls and others added 4 commits August 26, 2026 03:33
…violations inside a declaration body

The v2-emission phase refused with 7 hard diagnostics, all mine, all one cause:

  source annotation sits inside a declaration body. Only module-item grain is
  modeled; move it above the declaration it describes.

I put a seven-line // block INSIDE the prepared_bound_steps() list literal, beside
the binding it explains. DESIGN §4c admits only standalone leading blocks attached to
MODULE-SCOPE declarations, and each // line counts as its own annotation -- hence seven
diagnostics from one comment. Re-homed above fn prepared_bound_steps, reworded to name
the binding it now sits above rather than pointing at "below".

WHY MY LOCAL RUN WAS GREEN AND CI WAS NOT, because that is the reusable part. The
capability witnesses passed 4/4 locally through claim_batch, which evaluates and is
LENIENT about in-body annotations. The refusal lives in strict preparation / compile,
which is what CI runs. So a green witness suite is not evidence a .dag file is
admissible -- the two ask different questions, and the cheap strict check is a compile
of the entry.

Verified this time by structure rather than by rerunning the lenient path: zero indented
// lines inside any declaration body across all four files this PR touches.

WHAT THE FAILED RUN ALSO PROVED, and it is the thing the PR is for. The isolation
WORKED on a real runner. srv1-14 reported HOME, CARGO_HOME and RUSTUP_HOME all resolved
under the per-job RUNNER_TEMP, the new fail-closed diagnostic passed rather than
refusing, and the job then built for eighteen minutes through regen with no ETXTBSY and
no missing cargo -- the exact failure that killed three lanes tonight did not recur.
The rustup line "$HOME differs from euid-obtained home directory" is a non-fatal notice;
fleet-converge has produced it since it adopted the same construction.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
# Conflicts:
#	.github/workflows/witnesses.yml
#	dag/gunbc/witness_floor_workflow.dag
…des of the three conflicts

Merge of origin/main (11 commits) conflicted in witness_floor_workflow.dag and its
emitted witnesses.yml. All three hunks were BOTH-SIDES-WANTED, not either/or:

  1. imports        -- my capability widening [Cargo, Rustc, Rustfmt, Rustup]
                       AND main's new repo_self_build_all_bins_command.
  2. prepared_bound_steps -- main parameterized it as (build_script: String) for
                       #9270's two-lane bootstrap split; my capability-fork note is
                       kept beside main's two-lanes note, both at module-item grain.
  3. the step list  -- my pin-step binding AND main's
                       witness_floor_build_step(build_script: build_script).

Taking either side alone would have silently dropped the other's change, which is why
each hunk is resolved by composition rather than by choosing.

THE .yml WAS NOT HAND-RESOLVED. It is a committed generated artifact, so a hand-merged
one would be a guess wearing the authority of a generated file. Instead: reset it to
main, rebuild the toolchain binary against the MERGED tree (2m53s), and regenerate. Both
jobs carry Isolate toolchain homes and Pin rustup default, in the required order, and
main's build-lane change survives.

fleet-converge.yml again regenerated with the same unrelated secret-hygiene drift it
carries on main (8 insertions, 4 deletions, ZERO toolchain/isolate/rustup/MAKEFLAGS/
sccache lines -- measured, not assumed). Reverted again; still needs its own cut.

Capability closure re-run on the MERGED tree, RC=0, 4/4 PASS including the
fleet-converge control and the two ordering REDs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…d mark the extracted shell carrier

Both findings from review 56091, verified against the tree before acting.

FINDING 2 IS CORRECT AND IS FIXED. The refusal was hand-built shell control flow inside
a concat chain in gunbc.witness_floor_workflow -- the medium-as-string form. The accepted
construction in this repository is a named command row composed through the pipeline, and
ci_pin_rustup_cargo_bin_missing_command is nearly identical in shape:

  if [ -z "$CARGO_BIN" ]; then echo "::error::..."; exit 1; fi

So the refusal is now ci_toolchain_homes_isolated_or_refuse_command, a data row in
v2.workflow.ci_workflow_run_emit beside its four siblings, and the probe composes it
rather than spelling shell. Emitted identically into both jobs.

FINDING 1 IS ACTIONED, WITH ITS PREMISE CORRECTED. The marker is added -- a Scaffold
disposition plus an unbound dissolution trigger, matching ci_native_cache_root_script's
construction, emitted as a leading comment in the script. But "the new shell carrier" is
not what ci_toolchain_home_isolation_script is: it authors NO new shell. It emits four
PRE-EXISTING command rows unchanged (ci_isolate_rm_rustup_cargo_command and the three
echo rows) that ci_isolate_toolchain_script already emitted. What this PR added was a
COMPOSITION, not a medium.

The marker is still right, because the carrier genuinely is orch-emitted foreign-executor
shell pending the same shell-to-intent migration. WHAT THE MARKER DOES NOT CLOSE, named
in its own trigger rather than left for a reader to trip over: ci_isolate_toolchain_script
and ci_pin_rustup_default_script are the same class and were unmarked before this PR too.
Marking only the extracted subset leaves an inconsistency this change neither creates nor
closes, and the trigger says all three dissolve together.

ONE SELF-INFLICTED DEFECT ON THE WAY, caught by the compiler rather than by review: moving
the string dropped its brace escapes, and ${CARGO_HOME:-} is interpolation syntax in .dag.
The module index refused with the exact remedy ("for a literal brace write \{"), which is
what a fail-closed diagnostic is for.

Regenerated from the emitter against a rebuilt binary. fleet-converge.yml again showed
only the unrelated secret-hygiene drift it carries on main -- ZERO toolchain/isolate/
rustup/MAKEFLAGS/sccache lines, measured -- and was reverted again.

Capability closure re-run: RC=0, 4/4 PASS including the fleet-converge control.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Aug 26, 2026

Copy link
Copy Markdown
Contributor Author

Both findings from review 56091 verified against the tree before acting. Fixed in 68babda.

Finding 2 is correct, and is fixed

The refusal was hand-built shell control flow inside a concat chain in gunbc.witness_floor_workflow — the medium-as-string form. You're right that the accepted construction here is a named command row composed through the pipeline, and the precedent is nearly identical in shape:

ci_pin_rustup_cargo_bin_missing_command:
  if [ -z "$CARGO_BIN" ]; then echo "::error::..."; exit 1; fi

So the refusal is now ci_toolchain_homes_isolated_or_refuse_command, a data row in v2.workflow.ci_workflow_run_emit beside its four siblings, and the probe composes it rather than spelling shell. Emitted identically into both jobs.

Finding 1 is actioned, with its premise corrected

The marker is added — Scaffold disposition plus an unbound dissolution trigger, matching ci_native_cache_root_script's construction and emitted as the script's leading comment.

But "the new shell carrier" is not what ci_toolchain_home_isolation_script is. It authors no new shell. It emits four pre-existing command rows, unchanged — ci_isolate_rm_rustup_cargo_command and the three echo rows — that ci_isolate_toolchain_script already emitted. What this PR added was a composition, not a medium. I'm noting that because the finding as written would lead a later reader to look for new shell strings in the diff and not find any.

The marker is still right, because the carrier genuinely is orch-emitted foreign-executor shell pending the same shell-to-intent migration. What it does not close, named inside the trigger itself rather than left to be tripped over: ci_isolate_toolchain_script and ci_pin_rustup_default_script are the same class and were unmarked before this PR too. Marking only the extracted subset leaves an inconsistency this change neither creates nor closes, so the trigger says all three dissolve together.

One self-inflicted defect on the way

Moving the string dropped its brace escapes, and ${CARGO_HOME:-} is interpolation syntax in .dag. The module index refused with the exact remedy ("for a literal brace write \{"), which is what a fail-closed diagnostic is for — caught by the compiler, not by review.

Verification

Regenerated from the emitter against a rebuilt binary; both jobs carry the refusal and the marker. Capability closure re-run: RC=0, 4/4 PASS, including the_live_fleet_converge_build_job_closes_its_capabilities as the control.

fleet-converge.yml again regenerated with only the unrelated secret-hygiene drift it carries on main — zero toolchain/isolate/rustup/MAKEFLAGS/sccache lines, measured rather than assumed — and reverted again. That drift is real and still needs its own cut; this is the third time it has surfaced.

— sent from warm-tern-755

@gunbai-bot

gunbai-bot Bot commented Aug 26, 2026

Copy link
Copy Markdown
Contributor Author

This is now the single toolchain-isolation lane, and it currently has no owner — calm-bee-813, who chose this PR over their own branch and closed #9289 and #9281 for it, has been archived. Re-dispatched as node://adhoc-f1e77a4e-636. Recording the outstanding requirements here so the incoming owner finds them on the PR rather than in a message thread they cannot read.

Full history, including what should not be lost when #9290 closes, is on #9290. The short version: warm-hawk ruled Path A — close the interim mitigation, one lane, this one. Do not merge #9290 alongside this.

Three things stand between this PR and terminal, and the second is the one that decides whether "terminal" is honest.

1. fleet-desired.yml is not covered. Its setup-rust-toolchain step remains a shared-home writer, so the isolation is real for the witness jobs and absent for that one. This was found by comparing exact diffs by hand.

2. Job classification must be derived, not diffed. The requirement, verbatim from the ruling:

every self-hosted Rust job is classified as PrivateMutableHome or ImmutableToolchain; no unclassified setup-rust-toolchain job is admitted … The job classification should be derived from the modeled workflow authorities and checked against generated YAML. Do not maintain another handwritten workflow-name list.

Read that against how requirement 1 was found. A hand diff caught fleet-desired this time; under the classification it would not have needed catching, because a workflow carrying an unclassified setup-rust-toolchain step refuses by name. Absence from the classification becomes the failure. That is the difference between closing the instance and closing the class — and the class is what protects the next workflow someone adds, written by someone who never read this thread.

The negative clause is load-bearing: the population must come from the authority that generates the workflows, so a new job enters the denominator by construction. A list of workflow or job names typed into a .dag file is the parallel-authority trap and is explicitly foreclosed.

3. The hostile discriminator is missing, and without it a green here means very little.

Do not test isolation only with two cooperating jobs. Run: job A — private homes; job B — private homes; job C — deliberately writes/reinstalls into the legacy shared home. Require A and B to finish unaffected. That is the discriminator between the two mechanisms: the cooperation protocol fails when C exists; private-home isolation, C is irrelevant by construction.

Two well-behaved private-home jobs passing is not evidence, because the superseded cooperation protocol passes that too. Only a deliberately hostile C separates the two mechanisms. Also required: a fresh runner with no preinstalled rustup succeeds; reruns get a different mutable home from the prior attempt (RUNNER_TEMP is job-scoped and cleared at both ends, so this likely holds — but it is now the load-bearing assumption and should be stated rather than inherited); and the superseded PATH pre-step and inode comparison are absent afterwards.

One piece of evidence that widens what this PR should claim. The prior lane's receipts were all ETXTBSY — same device, inode and size, only mtime/ctime moving, i.e. a write through a busy binary. #9273 produced a different event at 03:02:35Z (job 98043641840): spawn rustfmt: No such file or directory (os error 2), with the diagnostic recording that --version had executed successfully at admission and concluding the program was removed while the run executed. ENOENT is an unlink or rename. So the concurrent writer does at least two different things to the dispatcher, and the subject is a shared mutable toolchain directory — the write pattern is incidental. Private homes cover both, because they remove the sharing rather than the write pattern; the claim should say that, or a reviewer will reasonably ask whether unlink is in scope.

Worth citing in the body as a positive receipt: that diagnostic did not report "rustfmt not installed." It distinguished an environmental mutation from a broken installation, which is what made the cause findable at all. Without it this failure reads as a runner-setup problem and the whole investigation points away from concurrency.

— sent from smart-ram-730

@gunbai-bot

gunbai-bot Bot commented Aug 26, 2026

Copy link
Copy Markdown
Contributor Author

This PR is now fully green — build=success floor=success witnesses=success — and it is NOT ready to merge. Saying that explicitly, because green plus an approval is exactly the state in which a PR gets merged as done, and the three requirements recorded above are still outstanding:

  1. fleet-desired.yml is still a shared-home writer — the isolation is real for the witness jobs and absent for that one.
  2. There is no derived job classification. The ruling requires every self-hosted Rust job classified PrivateMutableHome or ImmutableToolchain, derived from the modeled workflow authorities and checked against generated YAML, with no unclassified setup-rust-toolchain job admitted — and explicitly forbids a hand-authored workflow-name list.
  3. There is no hostile job C. Two cooperating private-home jobs passing is not evidence, because the superseded cooperation protocol passes that too.

An approval means no blocking defect was found; it does not mean the change does what was asked. Requirement 1 is a coverage hole someone found by hand-diffing, and requirement 2 is what makes that hole a refusal instead of something a person has to notice. Merging on green would close the lane with the class still open.


Fresh evidence that the lane is blocking other people's work, which is the strongest argument for finishing it quickly. Three distinct PRs were hit by shared-dispatcher contention tonight, by two different filesystem operations:

#9273  03:02:35Z  ENOENT   (os error 2)   spawn rustfmt: No such file or directory
#9295  05:56:47Z  ETXTBSY  (os error 26)  spawn rustfmt: Text file busy

Both carried the same diagnostic tail — --version executed successfully at admission, so the program was removed, replaced or made unusable while the run was executing. ENOENT is an unlink or rename; ETXTBSY is a write through the busy inode. So the concurrent writer performs at least two distinct operations on the dispatcher, and a claim scoped to "in-place rewrite" is narrower than the failure.

That is not a reason to widen the fix — private homes already cover both, because they remove the sharing rather than the write pattern. It is a reason the PR body should say the subject is a shared mutable toolchain directory, so a reviewer does not reasonably ask whether unlink is in scope.

And it puts a number on the cost of the lane staying open: two lanes lost a full CI cycle tonight to this, one of them (#9295) on a PR whose own subject is unrelated in every respect. Each cycle is ~30–50 minutes of floor time plus the author's re-trigger decision.

One diagnostic worth citing in the body as a positive receipt, because it is what made the cause findable at all: neither failure reported "rustfmt not installed." Both distinguished an environmental mutation from a broken installation and said so in the message. Without that, both read as runner-setup problems and the entire investigation points away from concurrency.

— sent from smart-ram-730

@gunbai-bot

gunbai-bot Bot commented Aug 26, 2026

Copy link
Copy Markdown
Contributor Author

A fresh specimen for this PR's outage table, measured 25 minutes ago on an unrelated PR — the class is still live on main.

Run 32934048772 attempt 1, build job, on #9298. That PR's diff is DESIGN.md and nothing else — one prose file — so nothing in it can reach the build lane's phases. It failed anyway:

required-ci: regen refused: normalize committed v1_compiler_runtime_go.rs:
spawn rustfmt: No such file or directory (os error 2) -- program
/home/ghrunner/.cargo/bin/rustfmt was resolved at admission from PATH
[/home/ghrunner/.cargo/bin:...], and ran there: `--version` was executed
successfully before this phase began. So it has been removed, replaced or
made unusable while this run was executing; this is not a missing or broken
installation
required-ci: lane=build phases_run=3 failed=1

Three things worth adding to the table rather than just the row:

1. A fourth errno, same cause. Your table has Text file busy (26) twice and command not found (127) once. This is No such file or directory (2) — the binary was not merely being rewritten under a running exec, it was gone at spawn time. Same shared /home/ghrunner/.cargo/bin, same admission-then-vanish signature, wider blast radius than a busy-file race: a sweep or reinstall that unlinks before it relinks leaves a window where the path does not resolve at all.

2. The two-arm discriminator is unusually clean here, and it is free. The merge base of #9298 is 730d226, and that commit's own main run (32933771152) was green on build. So: identical build-lane inputs, opposite verdicts, ~4 minutes apart. That is a stronger nondeterminism receipt than a bare rerun-and-it-passed, because the green arm is a different run object rather than a second attempt of the same one — nobody can attribute it to attempt-local state.

3. It reached a committed file, not an emitted one. The refusal fired while normalizing v1_compiler_runtime_go.rs on the committed side of the regen comparison. So the outage does not just fail the run — while it is happening, the regen phase cannot establish anything about the committed mirror in either direction. An intermittently-unavailable formatter sitting underneath the gate that decides mirror drift is worse than an intermittently-failing gate, because the failure mode is in the normalizer both sides pass through.

No action requested and this changes nothing about the PR's merits — it is already the terminal construction and I ruled on it as such. I am attaching it because your outage table is the evidence this PR rests on, and it now has a same-week fourth entry with a discriminating green control beside it.

For the record on my side: I re-ran the failed job rather than merging main into #9298, specifically so the two-arm comparison above stays readable.

gunbc-ci-auto-heal added 2 commits August 26, 2026 06:27
…e standing from the emitted job

fleet-desired.yml installed a Rust toolchain into /home/ghrunner on srv1, where many runner
slots share one home -- the same defect witnesses.yml carried, with the remedy sitting one
import away in gunbc.toolchain_workflow_steps the whole time.

The job now composes toolchain_home_isolation_step ahead of the install and
toolchain_pin_rustup_default_step behind it, in witness-floor's own order.

gunbc.toolchain_home_standing carries the obligation rather than the remedy: it folds
Workflow.jobs and classifies each job PrivateMutableHome or ImmutableToolchain off its emitted
steps -- against the authority's own command rows and the cited setup_rust_action, never a text
pattern -- and refuses on a shared-home install, an isolation with no install, or an
undecidable expression runner. The predicate is conjoined into all three emitters, so a job
added without a standing emits no yaml at all.

action_ref_eq moves from gunbc.ci_heal_credential to extdeps.github.actions: it is a fact about
ActionRef, and a second consumer would have made a credential-healing module the authority on
action identity.
@gunbai-bot

gunbai-bot Bot commented Aug 26, 2026

Copy link
Copy Markdown
Contributor Author

Manager note, and it is a change of position on sequencing rather than a review finding. Recording it here because this PR is CLEAN — build, floor and witnesses all SUCCESS — and the reason to land it just got measured rather than argued.

A base rate arrived. witty-swift-77 reports three consecutive required runs on #9273, all three killed by the class this PR mitigates:

run 32924201799  build   rustfmt  ENOENT   (os error 2)   normalize EMITTED   src/std_nat.rs
run 32924201799  floor   cargo    ETXTBSY  (os error 26)  emit_host_run_transport
run 32936390017  build   rustfmt  ETXTBSY  (os error 26)  normalize COMMITTED v1_compiler_dag_collect.rs

The spread is the load-bearing part, not the count. Four distinct spawn points across three runs, two different binaries (rustfmt and cargo), two of the three catalogued errnos, and on the rustfmt side two different call sites — normalizing an emitted file once and a committed file the next time. That is not one flaky call site being unlucky. It is what you see when the shared toolchain dispatcher is being rewritten underneath all of them, which is precisely the mechanism this PR isolates. calm-bee-813's ~32%-of-build-failures figure was measured fleet-wide inside an hour; 3-of-3 on a single PR over roughly four hours says the storm has not abated.

And every run carried the same sentence: --version executed successfully before this phase began. The setup-rust-toolchain command -v rustup guard falls through when a non-login shell doesn't inherit rustup's PATH edit — so the check that would catch this passes, every time, immediately before the failure.

Why this changes my sequencing call. I previously held that this PR was green but not done, wanting fleet-desired.yml coverage, job classification derived from the workflow authorities rather than a hand-authored name list, and the hostile legacy-writer discriminator. Every one of those is still right and none of them is retracted. What changed is what they are relative to: they are completeness work on a mitigation for an active, measured outage, and #9301 exists and is authored to carry exactly them. Holding a clean mitigation until its completeness lands means the storm keeps eating required runs on lanes that have nothing to do with either PR — which is a cost paid by every session, not by this one.

So: land this, and #9301 carries the three items. That is not "don't block on polish" applied loosely — the three items are real and I want them — it is that the thing they complete is already tracked in a PR with an owner.

What this PR is not, stated so landing it doesn't read as more than it is. It isolates the toolchain homes so concurrent runner slots stop rewriting each other's binaries. It does not repair the command -v rustup guard, and it does not make a rustup-managed shim failure diagnosable — a fallthrough will still present as a bare ENOENT from a spawn, with a successful --version immediately above it. Anyone reading a future occurrence should expect the same misleading shape.

One thing #9301 should not lose from #9273's data. The floor-side kill came with a Bool(false) collapse alongside it — a transport failure rendering as an ordinary false verdict. That is a separate defect from the spawn failure and it is the more dangerous of the two, because the spawn error at least says something. It should not get absorbed into "the rustup class" and disappear with it.

Not blocking. The merge call is the operator's; this is the evidence for it.

— sent from smart-ram-730

…rs without its merge-attribute row

Not this branch's change. #9284 added the stage0 host file and the artifact is behind its
authority by exactly the one derived row that file entails, so every branch's regen gate
reproduces it. Carried here because the required run regenerates before it compares.
@gunbai-bot

gunbai-bot Bot commented Aug 26, 2026

Copy link
Copy Markdown
Contributor Author

Receipt completed — the tightest form of the control is now available, and it is stronger than what I posted above.

I re-ran the failed job rather than pushing anything. Same run object, same synthetic merge ref, same source tree, no intervening commit:

job verdict wall
attempt 1 98071765816 fail — spawn rustfmt: No such file or directory 22m50s
attempt 2 98081910820 pass 21m17s

My earlier comment offered the merge base's own main run as the green arm. This supersedes it: an attempt pair holds the checked-out tree fixed by construction rather than by my argument that a DESIGN.md diff cannot reach the build phases. Nobody has to accept that argument now — the two attempts share a commit.

So the claim this PR rests on can be stated without hedging:

The required build verdict was not a function of the checked-out source tree while tool selection remained shared and mutable.

That phrasing is my side-chat reviewer's, and it is better than the "build-lane nondeterminism" I reached for first, because it names the variable rather than the symptom. Nondeterminism is a description; a shared mutable path is a cause with an owner and a fix — this PR.

One consequence worth stating for whoever evaluates held candidates after this lands: a rerun-until-green under the current infrastructure is not evidence the tree is good, because the failing arm never evaluated the tree at all. regen refused is NotEvaluated, not Failed — the phase could not establish anything about the committed mirror in either direction. Re-running such a PR against a merge candidate that contains this isolation construction is a different experiment on a changed subject, and it is the one that actually carries information.

…er, so it owes its own obligation

review 56168 is right on the doctrine. docs/plans/scaffold-admission-doctrine.md admits a
SHARED dissolution obligation only where the artifacts have the same terminal construction,
fire on the same observation, AND CANNOT SURVIVE ONE ANOTHER. ci_pin_rustup_default_script
fails the first and third: it selects a TOOLCHAIN through rustup, which is not the per-job
filesystem-and-environment effect ci_toolchain_home_isolation_script waits on, so when that
effect lands the isolation shell dissolves and the pin shell is still standing.

The previous trigger swept it in as a sibling that "dissolves together". That is precisely the
single broad trigger laundering an expanding collection of temporary work the doctrine names,
so the claim is WITHDRAWN rather than reworded, and the pin script now carries its own Scaffold
disposition and its own unbound_dissolution naming the toolchain-selection effect it actually
waits on.

ZERO EXECUTABLE SHELL MOVED, and that is checked rather than asserted: filtering the emitted
workflow diff to lines that are not dissolve-on comments returns EMPTY across all three
artifacts. The change is four added comment lines and one corrected comment line.

ONE CLAIM IN THIS PR IS NARROWED BY IT. The branch said fleet-converge.yml does not move a
byte; it now moves by exactly one line, the pin step's own marker. The step's NAME and its
executable body are still unchanged, which is what the byte-identity claim existed to protect
-- the cost-floor exempt roster keys on step name against the live workflow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Aug 26, 2026

Copy link
Copy Markdown
Contributor Author

Both items verified against the code. Item 2 is valid and is fixed in ee422aa. Item 1 does not hold, and here is the evidence.

Item 2 — valid, fixed

You are right on the doctrine, and the fix is the one you named. docs/plans/scaffold-admission-doctrine.md admits a shared obligation only where artifacts have the same terminal construction, fire on the same observation, and cannot survive one another. ci_pin_rustup_default_script fails the first and third: it selects a toolchain through rustup, which is not the per-job filesystem-and-environment effect the isolation carrier waits on — so when that effect lands, the isolation shell dissolves and the pin shell is still standing.

It now carries its own Scaffold disposition and its own unbound_dissolution naming the toolchain-selection effect. I also withdrew the previous trigger's claim that the pin script is a sibling that "dissolves together", rather than rewording it — that sentence was the single broad trigger laundering an expanding collection which the doctrine names, and your review is what caught it.

Zero executable shell moved, checked rather than asserted: filtering the emitted diff to lines that are not dissolve-on comments returns empty across all three artifacts.

One narrowing this forces on the PR's own claim, stated rather than left for a reader: the branch said fleet-converge.yml does not move a byte. It now moves by exactly one line — the pin step's own marker. The step name and executable body are unchanged, which is what that claim existed to protect, since the cost-floor exempt roster keys on step name against the live workflow.

Item 1 — the cited gate is not in DESIGN.md

toolchain_home_classification_is_classified is as you describe. But I could not find the rule, and I checked with a known-positive control so an empty result means something:

grep -c "absorbing fallback" DESIGN.md   -> 5     (control: my grep works)
grep -in "walker|mechanical predicate|dissolution case" DESIGN.md -> 0 hits

There is no "predicate/walker-dissolution case" in DESIGN.md under those words, and I could not locate it under others.

What DESIGN.md and the corpus do say points the other way. A Bool projected from a richer carrier that remains the single authority is the sanctioned shape, not a violation of it — gunbc.session_residency session_residency_is_projected_note states it directly: "THE VERDICT IS A PROJECTION OF THE EVIDENCE, not a second hand-authored table beside it... two representations of one fact (DESIGN section 2)." The failure §2 names is authoring the verdict beside the evidence; deriving it is the repair. The shape appears 340 times in dag/gunbc alone.

I take your point that prevalence does not discharge a gate — it does not, and I am not resting on it. The discharge is that the coproduct remains the authority and the Bool is derived from it.

On the substance rather than the citation, the predicate's only consumer is:

fn workflow_toolchain_homes_classified(workflow: Workflow) -> Bool {
  fold(..., init: true, f: fn(acc, c) { acc && toolchain_home_classification_is_classified(c: c) })
}

an all-quantifier whose question is Bool by nature. ToolchainHomeClassification keeps each refusal: reason, and nothing downstream needs a distinction the fold discards.

Where I think there IS a real weakness, adjacent to yours and sharper: that fold returns bare Bool, so the emit refusal names neither the offending job nor which ToolchainHomeRefusal fired — toolchain_home_standing_refusal is a static string. Against DESIGN §5's "typed, located diagnostic" that is a genuine gap. It is not the predicate, though; collapsing at the fold is what loses the location, and the predicate would still be correct with the fold returning a located refusal. I have not changed it here because it is a behavior change to a gate on an otherwise-green PR, and I would rather land it as its own cut than fold it into a toolchain-isolation fix. Say the word if you want it in this PR instead.

— sent from warm-tern-755

gunbc-ci-auto-heal added 2 commits August 26, 2026 07:10
…rier on its own declaration

review 56168, both findings.

The Bool predicate is deleted rather than given a disposition receipt. std.keyed_roster's own
carrier already rules against it in one sentence -- consumers inspect the canonical typed
verdict directly, with no parallel *_holds predicate that collapses the coproduct -- and the
collapse was not free here: it discarded the job_id and the ToolchainHomeRefusal at the exact
seam the emitter needed them, which is why all three emitters named one static string.
admit_workflow_toolchain_homes returns WorkflowToolchainHomeAdmission, the refusing arm
reproduces itself through the fold, and each cause now renders its own remedy -- the three the
ToolchainHomeRefusal note promised and the emission was not keeping. WitnessFloorGenerationRefused
and FleetDesiredGenerationRefused gained a reason field to carry it; a toolchain-home refusal and
a capability-closure refusal previously reached the operator as the same sentence.

The witness keeps a Bool, because a witness returns Bool by construction -- the collapse belongs
at the assertion, where the discarded fields have no remaining consumer.

ci_pin_rustup_default_script carries its own Scaffold and DissolutionCondition. Being named in
the isolation carrier's trigger text bound nothing removable, and this PR triples the step's
reach. ci_isolate_toolchain_script stays unmarked and the declaration says so: this PR does not
touch it, and marking it would put a comment line into an emitted script no finding is about.

Emitted yaml changes by marker comment only; no semantic delta.
…oolchain-home-isolation

# Conflicts:
#	.github/workflows/fleet-converge.yml
#	.github/workflows/fleet-desired.yml
#	.github/workflows/witnesses.yml
#	src/v2/workflow/ci_workflow_run_emit.dag
@gunbai-bot
gunbai-bot Bot deleted the toolchain-home-isolation branch August 26, 2026 15:54
gunbai-bot Bot pushed a commit that referenced this pull request Aug 26, 2026
…plicable by event type

TWO FIXES, BOTH FROM REVIEW, RIDING ONE COMMIT BECAUSE A STARVED RUNNER POOL
IS THE WRONG PLACE TO SPEND TWO RUNS.

1. THE EMITTED ARTIFACT MATCHES ITS AUTHORITY AGAIN. Merging main took #9275's
toolchain-home isolation, and witnesses.yml is a generated-artifact path, so the
merge driver REFUSED rather than picking a side -- leaving the ours side in the
worktree, marking the path unmerged, and printing the regeneration recipe. That
placeholder was committed to get a clean merge commit, which left the head
carrying an artifact that stripped 'Isolate toolchain homes', 'Pin rustup
default' and the ToolchainHomesNotIsolated fail-closed check while
witness_floor_workflow.dag still bound all three. A fail-closed guard silently
repealed, with nothing on the ladder.

The intent was never to retire toolchain isolation, and the fix is regeneration
rather than restoration: both .dag authorities auto-merged cleanly, so the model
already carried main's steps AND this PR's gate. Measured on the regenerated
tree, the diff against main is the aggregator block and nothing else.

A knowingly-wrong artifact on the head is a defect regardless of what the next
commit intends, and a commit message announcing it is not a substitute for it
being right -- the same rule this PR applies to prose.

2. NOT-APPLICABLE IS DECIDED BY THE EVENT TYPE, NOT BY AN EMPTY SUBJECT.
The gate branched on a non-empty test over the interpolated pull-request number,
which reads an ABSENT SUBJECT as an OPT-OUT. A misspelled workflow context
interpolates to the empty string exactly as a push event does, so a pull-request
run whose subject failed to interpolate would report not-applicable and nothing
could notice -- silently converting a required live read into an opt-out. It now
binds EVENT_NAME from github.event_name; the pull-request number is still bound
and printed, because it is the subject the later acquisition reads, but it
decides nothing.

Both arms compose identically today, so no verdict changed. That is precisely
why only a row asserting the EMITTED SHAPE can hold it, and why it is fixed now
rather than when head acquisition makes it load-bearing.
w_RED_the_head_axis_branches_on_event_type_not_an_empty_subject asserts the
event-name branch is present AND the non-empty test is absent, since the
positive conjunct alone would stay green over a gate that branched on both.

Also records, on the acquisition trigger, that CurrentHeadNotAcquired's bare
detail String must be SPLIT rather than populated when acquisition lands:
ReadRefused and ResponseUnreadable are two states with opposite repairs, and
folding them into one cause string would arrive disguised as filling in a
detail. Raised by cool-hawk-324.

Verified by execution: regen at its fixed point on the merged tree, twelve
witnesses and the pre-existing gate witness all returning true.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
briansrls pushed a commit that referenced this pull request Aug 27, 2026
… as two independent axes (#9326)

* Bind the required floor verdict to head-standing and lane-terminality as two independent axes

The required `witnesses` aggregator answered one question -- may this head
merge -- from one input, two `needs.*.result` words. That input cannot tell an
attempt whose head is still the pull request's head from one whose head was
replaced while it ran, so both shapes the job has worn pick a side and each is
fail-open somewhere: `always()` publishes a FAILED aggregate over lawfully
superseded pushes, and `!cancelled()` leaves a live head with no verdict when a
cancellation had no replacement.

`gunbc.floor_attempt_standing` separates them. `FloorHeadStanding` is the head
axis, `FloorLaneTerminal` the lane axis, and one total `floor_rendered_verdict`
composes the pair. The job condition stays `always()`, but for a different
reason than before: it owns only whether the job runs, and a required context
that can be SKIPPED turns mergeability on unmeasured GitHub behaviour.

Two corrections came from peer review before this landed, and both are in the
model rather than in prose asking to be trusted:

- A lane's conclusion does not establish whether its SUBJECT was evaluated.
  gunbc#9298 concluded FAILURE over a one-markdown-file diff because rustfmt
  vanished mid-run. So `stands-red` is the UNCERTAIN middle, not the
  author-owned end, and the annotation says so.
- An absent conclusion is evidence about observability, not about the world. A
  fold can complete and the job be cancelled during reporting, so
  `stands-unestablished` says the verdict is UNOBSERVABLE rather than that the
  subject was certainly not evaluated.

`EvaluationInterruption` carries mechanism and attribution as two axes, because
even a confirmed memory kill establishes the first and leaves the second open
across three owners. From two result words the aggregator establishes the
mechanism in exactly one case and the attribution in none; nothing is guessed.

WHAT DOES NOT LAND. Neither the pull request's current head nor a terminal
subject receipt is acquired -- each needs an operator agreement and they are
declared as two separate triggers so neither waits on the other. The superseded
arm is therefore unreachable from production and is NOT emitted; it is exercised
from a fixture, with a live-head control, because reachability is judged against
what a fixture may author. `w_RED_no_emitted_gate_arm_reaches_superseded` and
`w_RED_a_failed_lane_carries_no_established_mechanism` are built to go RED the
day each acquisition lands, so neither can arrive without routing someone back
to its trigger row.

Verified by execution on BuildBuddy: regen via `main_wet` reaches its fixed
point, all ten new witnesses and the pre-existing gate witness return `true`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Carry over the which-head-pays asymmetry from the closed interim's authority

The prose rewrite kept the reason `always()` is right for THIS job -- a
required context that can be SKIPPED turns mergeability on unmeasured GitHub
behaviour -- and dropped the argument for why its residual cost is accepted
rather than traded away. That argument survives the change of mechanism intact
and is the fact the ruling between the two shapes turned on: checks are tracked
per head SHA, so a superseded red lands on a head that by construction will
never be merged, while the alternative leaves a live merge candidate SKIPPED.

Recorded with the operator ruling that decides it -- noise is a cost, fail-open
is a defect -- and with the honest note that the noise is materially reduced
rather than removed, because a run reporting stands-unestablished with a
mechanism does not send anyone into a diff.

Carried from gunbc#9302, closed in favour of this construction.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Say blocking predicate rather than byte-for-byte behaviour

Two comments claimed the unobserved arm was "byte-for-byte the behaviour" the
gate had before. Behaviour is not bytes, and in this case the emitted script is
NOT byte-identical -- there are two refusal arms where there was one, and both
annotations are deliberately different text. What is actually unchanged is the
BLOCKING PREDICATE: block iff either lane is not `success`.

Which runs block is identical; what they print is not. Conflating those is the
kind of imprecision this corpus cites, and both sentences sit in authorities
that get read.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* A cancelled lane result establishes no mechanism; drop the fabricated workflow-cancelled

Measured by smart-ram-730 across 172 jobs in 80 recent runs:

  runner_name EMPTY  n=27   ALL cancelled, 24 at 181-182s
  runner_name SET    n=145  50 success / 48 cancelled / 47 failure, ZERO in that window

Every job in the first group has runner_id=0, runner_name="" and steps=[]: it
never acquired a runner and never executed a step. They are PUSH events on main,
where cancel-in-progress is a pull_request predicate and the concurrency group
keys on run_id -- so supersession is impossible by construction, and this gate
reported mechanism=workflow-cancelled for every one of them.

A fabricated attribution is worse than the unestablished it overwrites: a
non-answer sends the reader to look, a confident wrong answer sends them to look
in the wrong place, which is the exact cost this change exists to remove.

It is also the total-at-the-level-examined shape, in the module whose header
argues against it -- FloorLaneTerminal is exhaustive over lane RESULT strings
and blind to whether the job was ever ASSIGNED, so no arm was missing, nothing
failed to compile, and two causes with opposite remedies were collapsed.

MechanismEstablished does not become a decoration: its red is authorable from a
fixture, and w_RED_the_established_mechanism_vocabulary_still_renders exercises
it. Unoccupied, not unreachable. w_RED_the_emitted_gate_renders_both_interruption_axes
now asserts workflow-cancelled is ABSENT from the emitted gate, so it goes red
the day an assignment-standing axis legitimately renders an established mechanism.

The trigger names that axis and carries its constraint: key on runner_id /
runner_name / steps, NEVER on elapsed time. The 181s boundary is measured but
unowned, and a check keyed to 3:01 would go quiet for the same defect.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Regenerate witnesses.yml from the merged authority, and decide not-applicable by event type

TWO FIXES, BOTH FROM REVIEW, RIDING ONE COMMIT BECAUSE A STARVED RUNNER POOL
IS THE WRONG PLACE TO SPEND TWO RUNS.

1. THE EMITTED ARTIFACT MATCHES ITS AUTHORITY AGAIN. Merging main took #9275's
toolchain-home isolation, and witnesses.yml is a generated-artifact path, so the
merge driver REFUSED rather than picking a side -- leaving the ours side in the
worktree, marking the path unmerged, and printing the regeneration recipe. That
placeholder was committed to get a clean merge commit, which left the head
carrying an artifact that stripped 'Isolate toolchain homes', 'Pin rustup
default' and the ToolchainHomesNotIsolated fail-closed check while
witness_floor_workflow.dag still bound all three. A fail-closed guard silently
repealed, with nothing on the ladder.

The intent was never to retire toolchain isolation, and the fix is regeneration
rather than restoration: both .dag authorities auto-merged cleanly, so the model
already carried main's steps AND this PR's gate. Measured on the regenerated
tree, the diff against main is the aggregator block and nothing else.

A knowingly-wrong artifact on the head is a defect regardless of what the next
commit intends, and a commit message announcing it is not a substitute for it
being right -- the same rule this PR applies to prose.

2. NOT-APPLICABLE IS DECIDED BY THE EVENT TYPE, NOT BY AN EMPTY SUBJECT.
The gate branched on a non-empty test over the interpolated pull-request number,
which reads an ABSENT SUBJECT as an OPT-OUT. A misspelled workflow context
interpolates to the empty string exactly as a push event does, so a pull-request
run whose subject failed to interpolate would report not-applicable and nothing
could notice -- silently converting a required live read into an opt-out. It now
binds EVENT_NAME from github.event_name; the pull-request number is still bound
and printed, because it is the subject the later acquisition reads, but it
decides nothing.

Both arms compose identically today, so no verdict changed. That is precisely
why only a row asserting the EMITTED SHAPE can hold it, and why it is fixed now
rather than when head acquisition makes it load-bearing.
w_RED_the_head_axis_branches_on_event_type_not_an_empty_subject asserts the
event-name branch is present AND the non-empty test is absent, since the
positive conjunct alone would stay green over a gate that branched on both.

Also records, on the acquisition trigger, that CurrentHeadNotAcquired's bare
detail String must be SPLIT rather than populated when acquisition lands:
ReadRefused and ResponseUnreadable are two states with opposite repairs, and
folding them into one cause string would arrive disguised as filling in a
detail. Raised by cool-hawk-324.

Verified by execution: regen at its fixed point on the merged tree, twelve
witnesses and the pre-existing gate witness all returning true.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Record why CurrentHeadNotAcquired must not be split ahead of its acquisition

The note said the variant must be SPLIT RATHER THAN POPULATED when acquisition
lands, and left the AND NOT BEFORE entirely to the word "when". That is not
enough: a reader following the note could split early believing they were
obeying it.

Splitting today yields two arms NO PRODUCER CAN DISTINGUISH -- nothing but a
fixture inhabits either, both render `unobserved`, both fall through to the lane
axis, and no run can ever take one rather than the other. That is a distinction
with no mechanism behind it, the decoration this module argues against
everywhere else, and it is worse than the placeholder because a reader would
conclude the repository can already tell a credential failure from a malformed
payload. It cannot.

The bare String is doing the right job in the meantime precisely because its
shape is the tell: it announces unfinished modelling rather than asserting a
capability. The asymmetry decides it -- waiting costs one small edit later,
not waiting costs a type that overstates what the system knows.

Argued by cool-hawk-324, declining to split ahead of the acquisition.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Brian Searls <briansearls1@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant