Repository navigation
spark grants: carry the subject in the variant, derive both renderings from one upstream binary authority - #8860
Conversation
…s from one upstream binary authority
The linger grant had FOUR spellings across the corpus and the two in
managed_access_bootstrap were both wrong against the live hosts.
MEASURED on srv5 and srv6, 2026-08-22, via `sudo -n -l` as the grantee:
(root) NOPASSWD: /usr/bin/loginctl enable-linger gunbc-automation
(root) NOPASSWD: /usr/bin/systemctl reboot
sudo -n -l loginctl enable-linger -> REFUSED
sudo -n -l loginctl enable-linger gunbc-automation -> ALLOWED
sudo -n -l systemctl reboot -> ALLOWED
spark_managed_grant_sudoers_command emitted "/usr/bin/loginctl enable-linger"
and spark_managed_grant_argv emitted bare ["loginctl", "enable-linger"], six
lines apart, and NEITHER carried the subject account. So the argv this module
derives cannot execute under the grant the hosts hold, and the grant this
module installs is narrower than the one the actuator needs.
serving_realization and serving_converge_realize already emit the correct
form, with the subject. They are untouched here: they are the Y this change
dissolves the wrong pair into, not a third thing to reconcile.
THE SUBJECT IS CARRIED IN THE VARIANT rather than threaded uniformly, because
the account axis is not uniform: enable-linger takes a subject, reboot takes
none. A grant authority parameterized by subject UNIFORMLY would authorize
`systemctl reboot <account>`, an argument sudo would refuse to match. Carrying
it per-variant makes the arity difference structural.
It also makes the variant's own name answerable. "ForOwnAccount" was a claim
in an identifier that nothing could check, because there was no account in the
value to compare the executor's login against. This change does not add that
check -- it stops the question being unaskable.
ONE AUTHORITY PER BINARY. extdeps.systemd gains systemd_binary_directory and
systemd_binary_path; extdeps.systemd.loginctl is new, cited to freedesktop's
loginctl(1); systemctl gains its name/path/reboot rows. The absolute rendering
(for the sudoers Cmnd_Spec) and the bare rendering (for the argv, resolved by
secure_path) now fold the same rows. They are one binary rendered for two
consumers, not two facts.
Bytes: the reboot command and both reboot argvs are UNCHANGED. The linger
command and argv MOVE, by exactly the subject token -- toward what the hosts
hold, not away from it.
THE TEST THIS REPLACES WAS VACUOUS AND ITS NAME SAID OTHERWISE.
the_sudoers_line_authorizes_exactly_the_argv_the_actuator_runs asserted
any(argv.skip(2), w => string_contains(command, w)) -- an ANY over substring
containment, satisfied by "loginctl" alone. It would have passed with the argv
truncated, with a bogus word appended, and did pass with the subject missing
from both sides while the hosts required it.
Replaced with an argv-grain equality plus two discriminating arms: the
pre-change spelling must not equal what is derived, and the derived argv must
equal the full five-word form. The two fixtures moved too: they had been
written from the model's own sudoers_command rather than captured from a host,
so they agreed with the defect by construction.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Evidence: green by execution, plus a discriminating RED by mutationEvery row is
All three new assertions red under the mutation. That is the point of running it: the test this PR replaces was an The five pre-existing reconcile witnesses stay green against the corrected fixtures — so moving Also verified: What is not claimedI have not run a full local floor. The evidence above is per-witness execution plus two clean compiles; the merge bar is the CI run on this head. One thing worth flagging for whoever merges: main is currently red at — sent from fierce-lynx-647 |
The failing check is main's, and the run additionally confirms this PR's witnesses are enrolled and greenRun 32546817865 @ Failure identities diffed against main's run at Identical. All 16 are Two things this run proves that my local evidence could not: 1. The new witnesses are enrolled. #8857 — a one-file change adding no tests, same base — planned 2. They pass in CI, not just locally. Also clean:
No fix pushed — there is no defect here to fix. This PR is unblocked when main is. — sent from fierce-lynx-647 |
Re-verified on the merged tree (
|
| test | merged (ships) | merged + subject dropped |
|---|---|---|
the_sudoers_command_equals_the_line_the_hosts_actually_hold |
true |
false |
the_sudoers_line_authorizes_exactly_the_argv_the_actuator_runs |
true |
false |
the_pre_change_spelling_without_the_subject_is_refused |
true |
false |
the_derived_argv_carries_the_subject_the_grant_requires |
true |
— |
Both arms complete: 4/4 green on what ships, 3/3 red under the mutation. The green arm is reported alongside the red deliberately — reds alone are equally consistent with having broken the subject and rewritten the tests to match.
The bytes are unchanged from the pre-merge verification: reboot command and both reboot argvs identical, the two linger strings carrying the subject token that srv5 and srv6 actually grant.
— sent from fierce-lynx-647
Correction to my earlier comment: the
|
| test | merged (ships) | + subject dropped |
|---|---|---|
the_sudoers_command_equals_the_line_the_hosts_actually_hold |
true |
false |
the_sudoers_line_authorizes_exactly_the_argv_the_actuator_runs |
true |
false |
the_pre_change_spelling_without_the_subject_is_refused |
true |
false |
the_derived_argv_carries_the_subject_the_grant_requires |
true |
— |
By CI: failed=16, an identity set identical to main's, so none of these three is among the failures. That is a set comparison, not a count, and it holds.
Not established: that these three executed in CI. Absence from the failure set is consistent with passing and also consistent with never having run — and per §5 an absence is not evidence. The module sits in dag/test/claim/, which discovery scans, and its sibling witnesses predate this PR, so non-discovery is unlikely; unlikely is not established, and I am not going to write it as though it were.
The identity diff I used for failures throughout is unaffected — comparing failing identities against main is a set comparison. That is precisely why it survived while the counts around it did not.
— sent from fierce-lynx-647
Tightening the correction above: discovery is deterministic, not driftingMy previous comment said discovery "has been measured moving with no corpus change." That wording was wrong twice: there was a change between those two commits (one line), and it implied nondeterminism. Nondeterminism is now refuted, and the controlled pair is on this very PR. This run had two attempts because its first failed on the rustfmt runner: Same commit, different runners, byte-identical figures. Discovery is a deterministic function of the commit, and stable across runners. The conclusion is unchanged; the reason is better
A confounded measurement and a noisy one fail identically for this purpose — neither isolates my change — but they are different facts, and the weaker claim is the true one. Reproducibility is real: a delta between a head and its base does measure something, just not only the witnesses added. And the identity replacement remains unavailableRestating the claim as "the three appear by name in the floor's executed pass list" is the right instinct and that list does not exist. On this run: passing witnesses emit no line (only So the honest position stands exactly as posted above: local execution with a discriminating mutation RED is the evidence; CI establishes only that these three are not among the failures. Writing the pass-list citation would have swapped a false inference for a false citation, which is worse. — sent from fierce-lynx-647 |
…at eats its own evidence TWO CHANGES ON TOP OF THE APPROVED SINGLE-FILE UPLOAD, both directed by the coordinating session. ALL THREE, NOT ONE. The floor writes three files on every run and discards all three: cli_run.rs reads GUNBC_REQUIRED_FLOOR_DISPOSITION, GUNBC_EXPECTED_RED_ROSTER_JOIN and GUNBC_LONG_HOME_STORAGE_AGREEMENT, all three are already set on this job, and witnesses.yml had zero upload-artifact steps. That is ONE unwired last hop with three instances, not three similar things, so it is fixed at the class by one step shape applied three times. Fixing one site would leave the other two for someone to rediscover, paying the whole investigation again to reach a conclusion already written down. The single-authority path move is now load-bearing three times: six literals would otherwise be free to drift into the worst failure shape, where if-no-files-found fires and reads as "the floor produced no roster" when the truth is "the writer and the uploader disagree about a filename". `if-no-files-found: error` was CHECKED, not assumed -- `error` on a file a run does not always write would redden green runs. The join file's write is guarded on roster_join_active, which is GUNBC_EXPECTED_RED_ROSTER_JOIN.is_some(), set here; the other two are guarded only on their own paths, likewise set. AN INSTRUMENT FOR THE rustfmt ENOENT. `spawn rustfmt: No such file or directory` refused a required run today, and the message reports the syscall and nothing about the environment that produced it. Two sessions produced three mutually incompatible readings from it in one night -- a per-host fact, a runner lottery, a missing component -- and every one was refuted by data the reading did not have. The rustup toolchain logged `component rustfmt is up to date` in the same job that could not spawn it, so a bare PATH lookup found neither a shim nor a toolchain binary, and no run records which exists. The new step records hostname/RUNNER_NAME, PATH, which -a rustfmt, which -a cargo rustc rustup, the cargo bin listing, rustup which rustfmt, rustup toolchain list and rustup default. It runs on GREEN jobs too, deliberately: nobody knows what a working job's PATH looks like either, so a failed run's environment could not be compared against anything. A zero is readable only beside a nonzero, and today there is neither. Guarded on `!cancelled()` alone -- weaker than witness_floor_precondition, which also requires the build to have succeeded. A build failure is precisely when the environment is in question, so gating the recorder on the build would blind it exactly when it is wanted. EVERY PROBE IS TOTAL BY CONSTRUCTION, so the step cannot fail and needs no continue-on-error. Not cosmetic: continue-on-error is a step that MAY report a verdict and has been told to be ignored -- the escape-hatch shape DESIGN 5 forbids. A step with no verdict to give is a different object. It also means a tool's ABSENCE records as a printed line rather than a dead step. WHAT THIS DOES NOT ANSWER, unchanged: a row reading Planned means DISCOVERED AND ROUTED. Not executed, not passed. ClaimOutcome::Pass discards the identity at accumulation, so no execution roster exists to upload. Narrows #8860; does not close it. A PREDICTION, so nobody spends an investigation on it: on a regen refusal the floor still RUNS -- the three phases are independent (run 32556977241: phases_run=3 failed=1 with planned=10439 executed=10439 passed=10132 after the rustfmt refusal) -- so the files are written and the uploads publish the roster of a REFUSED run, which is when it is most wanted. The case that does red all three is a FLOOR phase refusing before its writes, on an already-stopped line. Also adds a one-line carrier note on witness_step_status_guard, whose name says guard and whose value is a prefix ending in `&& `: used bare it emits a YAML condition ending in `&&`. Verified: regen exit=0 and clean, and workflow_capability_closure_witness the_live_witness_floor_job_closes_its_capabilities returns true with all four new steps enrolled. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
MEASURED, read-only `sudo -n -l` on both units 2026-08-22:
192.168.1.222 spark-a3ee (root) NOPASSWD: /usr/bin/loginctl enable-linger gunbc-automation
(root) NOPASSWD: /usr/bin/systemctl reboot
192.168.1.223 spark-3bd5 (same two)
ONE COMMAND PER LINE, no comma list -- each grant is its own drop-in and so its
own Cmnd_Spec. The comma shape does not occur on srv5 or srv6 today.
THE JOIN SPLITS ON COMMAS ANYWAY, because sudo renders a multi-command Cmnd_Spec
comma-separated on one line, and reading that line as ONE command would match
neither grant it names. That is a LIVENESS hole rather than a safety one --
convergence would never complete on a host that plainly holds both, and the
diagnostic would report grants missing that are visibly present. One line of
parsing, so it is taken rather than declared as a limit. And the split cannot
make a wrong answer right-looking: a comma inside a command's own ARGUMENT makes
fragments that match nothing, the grant reads not-held, the installer reinstalls
-- the same safe direction as not splitting, over a strictly larger set of
listings read correctly.
`NOPASSWD: ALL` READS NOT HELD, now named in the carrier before someone files it.
That is correct and fail-closed: teaching the join that ALL subsumes everything
would let a blanket grant answer for every specific one, which is the
host-grained "is provisioned" Bool this module exists to avoid.
THE COMMA TEST'S FIRST DRAFT WAS WRONG AND THE RUN CAUGHT IT. I hand-spelled the
fixture with the post-#8860 linger command, and it returned FALSE -- reporting a
comma-splitting defect when what it had measured was the subject's absence from
today's modeled command. Two ordering-dependent tests where one was intended, and
the second one lying about which axis it failed on. The fixture is now RENDERED
from spark_managed_grant_sudoers_command over the population, so it agrees with
the model on spelling by construction -- right here precisely because spelling is
the ordering witness's subject, not this test's.
BOTH NEW TESTS RE-VERIFIED AGAINST THEIR CONTROLS after this change, since a fix
can quietly turn a discriminator vacuous:
two_commands_on_one_entry_line_are_two_grants
comma split present returned `true`
comma split removed (control) returned `false`
an_entry_that_merely_starts_with_the_desired_command_is_not_that_grant
still returned `true` after the comma change -- the split did not erode it
THE LIVE READ ALSO CONFIRMS #8860 FROM THE HOST SIDE: the installed linger
command carries the subject, byte-identical to the constant the ordering witness
pins. The witness therefore measures against a line read off a host, not against
a fixture that agrees with the model by construction.
srv5/srv6 stand 2/2 by live host reading. The INSTRUMENT does not compute that
verdict until #8860 and #8852 land.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…raint instead of noting it (#8882) * Join the grant listing entry-by-entry, and execute the ordering constraint instead of noting it THE SUBSTRING JOIN WAS PREFIX-BLIND. `spark_managed_grant_standing_from_outcome` asked `string_contains(listing, command)` over the whole `sudo -n -l` blob, which answers HELD for any installed command the desired one is a prefix of. So a host authorizing MORE than the population desires read as converged, with nothing naming the excess -- and the grain at which sudo's answer is meaningful is one ENTRY, so a substring test over the concatenation is a test at no grain at all. The join is now entry-grained, parsed at sudo's own seam: each authorization renders as ` (root) NOPASSWD: <command>`, so the command is what follows the last `NOPASSWD: ` on its line, trimmed. Splitting on the marker rather than on whitespace keeps a command containing spaces one command. A line with no marker yields NO command rather than a false one, so banners and blank separators contribute nothing to the join instead of contributing a negative vote. THE DISCRIMINATING RED, AND THE FACT THAT MY FIRST ONE WAS VACUOUS. The new test feeds a listing whose entry ends in a subject the desired command does not name -- a wider grant, of which the desired command is a proper prefix. Measured, same test, two trees: entry-grained join (this branch) returned `true` substring join restored (control) returned `false` The first draft of that test asserted `!converged` and stayed GREEN under the mutation. It was measuring the reboot grant's absence from the fixture -- unheld under either join -- and nothing about the change it was written for. Counting the install set isolates the one grant the two joins disagree about: 2 under the entry-grained join, 1 under the substring join. The control is what caught it, which is the only reason the surviving test can be trusted. THE ORDERING WITNESS IS RED ON THIS BRANCH BY CONSTRUCTION. This fix and the grant-spelling correction in gunbc#8860 are two defects that CANCEL: the modeled linger command currently omits the subject the hosts hold, and the substring join is the only reason that mismatch does not surface. Tighten the join first and the linger grant reads NOT-HELD on two converged live hosts, and the reconcile reports a remedy to install a grant that is already there. A note in a PR body cannot enforce a merge order -- it depends on a human merging two PRs in one direction. So the constraint executes instead: `the_modeled_linger_command_is_byte_equal_to_the_line_the_hosts_hold` asserts the modeled command equals the line measured on srv5 and srv6. It returns `false` today and `true` once #8860 lands, so merging this branch first cannot go green. It reads only rendered command strings, never #8860's variant shape, so it compiles identically before and after that PR. DO NOT MERGE THIS BEFORE gunbc#8860. The witness says so by failing, not by asking. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Split multi-command entry lines, measured against both live hosts MEASURED, read-only `sudo -n -l` on both units 2026-08-22: 192.168.1.222 spark-a3ee (root) NOPASSWD: /usr/bin/loginctl enable-linger gunbc-automation (root) NOPASSWD: /usr/bin/systemctl reboot 192.168.1.223 spark-3bd5 (same two) ONE COMMAND PER LINE, no comma list -- each grant is its own drop-in and so its own Cmnd_Spec. The comma shape does not occur on srv5 or srv6 today. THE JOIN SPLITS ON COMMAS ANYWAY, because sudo renders a multi-command Cmnd_Spec comma-separated on one line, and reading that line as ONE command would match neither grant it names. That is a LIVENESS hole rather than a safety one -- convergence would never complete on a host that plainly holds both, and the diagnostic would report grants missing that are visibly present. One line of parsing, so it is taken rather than declared as a limit. And the split cannot make a wrong answer right-looking: a comma inside a command's own ARGUMENT makes fragments that match nothing, the grant reads not-held, the installer reinstalls -- the same safe direction as not splitting, over a strictly larger set of listings read correctly. `NOPASSWD: ALL` READS NOT HELD, now named in the carrier before someone files it. That is correct and fail-closed: teaching the join that ALL subsumes everything would let a blanket grant answer for every specific one, which is the host-grained "is provisioned" Bool this module exists to avoid. THE COMMA TEST'S FIRST DRAFT WAS WRONG AND THE RUN CAUGHT IT. I hand-spelled the fixture with the post-#8860 linger command, and it returned FALSE -- reporting a comma-splitting defect when what it had measured was the subject's absence from today's modeled command. Two ordering-dependent tests where one was intended, and the second one lying about which axis it failed on. The fixture is now RENDERED from spark_managed_grant_sudoers_command over the population, so it agrees with the model on spelling by construction -- right here precisely because spelling is the ordering witness's subject, not this test's. BOTH NEW TESTS RE-VERIFIED AGAINST THEIR CONTROLS after this change, since a fix can quietly turn a discriminator vacuous: two_commands_on_one_entry_line_are_two_grants comma split present returned `true` comma split removed (control) returned `false` an_entry_that_merely_starts_with_the_desired_command_is_not_that_grant still returned `true` after the comma change -- the split did not erode it THE LIVE READ ALSO CONFIRMS #8860 FROM THE HOST SIDE: the installed linger command carries the subject, byte-identical to the constant the ordering witness pins. The witness therefore measures against a line read off a host, not against a fixture that agrees with the model by construction. srv5/srv6 stand 2/2 by live host reading. The INSTRUMENT does not compute that verdict until #8860 and #8852 land. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Name the test that owns the spelling axis, so the rendered fixture is not "fixed" back The carrier already said the fixture renders from the model deliberately. It did not NAME the test that owns the axis it is deferring to, so a reader had the justification without the referent -- and the §3 rule is to cite the symbol. It now names `the_modeled_linger_command_is_byte_equal_to_the_line_the_hosts_hold` as the owner of the spelling axis, says why agreeing-by-construction is normally disqualifying, and says plainly not to convert this back to a hand-spelled literal -- which would silently restore the two-axis failure and make a parse test fail for a spelling reason. Re-verified after the edit: two_commands_on_one_entry_line_are_two_grants still returns `true`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
The defect
The linger grant had four spellings across the corpus. Two of them — both in
managed_access_bootstrap— were wrong against the live hosts.Measured on srv5 and srv6, 2026-08-22, via
sudo -n -las the grantee. No host was touched;sudo -lis a matchability query, not an execution:spark_managed_grant_sudoers_command/usr/bin/loginctl enable-lingerspark_managed_grant_argv[sudo, -n, loginctl, enable-linger]serving_realizationlinger arm/usr/bin/loginctl enable-linger <principal>spark_serving_enable_linger_remote_argv[sudo, -n, loginctl, enable-linger, <login>]3 and 4 match the hosts and are untouched here — they are the Y this dissolves 1 and 2 into, not a third thing to reconcile.
Why the wrong pair was not merely dead
Spelling 2 is unreachable in production today (the only non-test caller of
spark_managed_grant_argvpassesRebootThisHost, which is ALLOWED). But spelling 1 is the installer. Ifmanaged_grant_installran against a host it would write a grant narrower than the actuator needs, and linger convergence would break at the next apply.What this does not do, and the ordering that follows
Nothing here touches a host. Verified three ways:
fleet-converge.ymlisworkflow_dispatchonly; on main the step's entrydag/gunbc/spark/grant_install.dagdoes not exist; and install is gated on the measured-missing set.That last gate is the important one, because two defects are currently cancelling.
spark_managed_grants_to_installderives from the reconcile'sstring_contains(listing, command)— a substring join. The installed line is a superstring of spelling 1, so reconcile reportsGrantHeld, the install set is empty, and nothing is written. The substring bug is the only thing currently preventing the wrong installer from firing.So the order is forced, and the general rule is worth stating: when two defects cancel, fix the one being masked, never the mask. Fixing the reconcile join first would unmask an installer that writes a narrower grant onto two live hosts. This PR is the masked half. The reconcile fix follows, and will carry a witness that fails unless this has already landed — so the ordering is enforced by the tree rather than by merge discipline.
The modeling change
The subject is carried in the variant:
The account axis is not uniform —
enable-lingertakes a subject,reboottakes none. A grant authority parameterized by subject uniformly would authorizesystemctl reboot <account>, an argument sudo would refuse to match. Per-variant makes that arity difference structural.It also makes the variant's own name answerable: "ForOwnAccount" was a claim in an identifier with no account in the value to check it against. This PR does not add that check — it stops the question being unaskable.
One authority per binary.
extdeps.systemdgainssystemd_binary_directory/systemd_binary_path;extdeps.systemd.loginctlis new, cited to freedesktop'sloginctl(1);systemctlgains name/path/reboot rows. The absolute rendering (sudoersCmnd_Spec) and the bare rendering (argv, resolved viasecure_path) fold the same rows — one binary rendered for two consumers, not two facts.systemctl.dagstill spells"systemctl"inline at 24 further sites, mostly insidetransport shell { argv: [...] }DSL rows. Left alone deliberately and not claimed fixed — that is a separate, byte-neutral cleanup.Bytes
/usr/bin/systemctl rebootsudo -n systemctl reboot/usr/bin/loginctl enable-linger… enable-linger gunbc-automationsudo -n loginctl enable-linger… enable-linger gunbc-automationThe two linger strings move, by exactly the subject token — toward what the hosts already hold.
The test this replaces was vacuous, and its name said otherwise
An
anyover substring containment:"loginctl"is inside"/usr/bin/loginctl enable-linger", so it passed. It would have passed with the argv truncated to[sudo, -n, loginctl], with a bogus word appended — and did pass with the subject missing from both sides while the hosts required it. "Exactly" was a word in an identifier over a body asserting almost nothing.Both fixtures moved too:
srv6_listing_before/_afterwere written from the model's ownsudoers_commandrather than captured from a host, so they agreed with the defect by construction. They now carry the measured line.Evidence
Green by execution, plus a discriminating RED produced by mutation in an isolated worktree (subject dropped from
loginctl_enable_linger_command, i.e. the pre-change spelling restored) — see the table posted in the thread below.