diff --git a/DESIGN.md b/DESIGN.md index 59312f9ba9b..5352329d30d 100644 --- a/DESIGN.md +++ b/DESIGN.md @@ -239,6 +239,7 @@ One row per class, each carrying its recognition rule and its receipts, in [docs - `liveness_probe_read_as_currency` - `non_verdict_disposition_surfaces_as_refusal` - `empty_capture_read_as_clean_result` +- `salience_instrument_blind_to_the_record_it_sizes` - `ambient_process_state_read_by_a_concurrent_reader` - `predicate_vacuously_true_on_an_empty_domain` - `check_subject_narrower_than_its_declared_claim` diff --git a/dag/gunbc/recurring_failure_mode.dag b/dag/gunbc/recurring_failure_mode.dag index 1c9b40c07d4..c44db243d72 100644 --- a/dag/gunbc/recurring_failure_mode.dag +++ b/dag/gunbc/recurring_failure_mode.dag @@ -183,6 +183,8 @@ data sealing_property_erases_structure: RecurringFailureMode = RecurringFailureM data censored_estimator_drops_its_own_tail: RecurringFailureMode = RecurringFailureMode { identity: "censored_estimator_drops_its_own_tail" as NonEmptyStr, authored: "**censored estimator drops its own tail** (an estimate of a variable is computed over the observations that SURVIVED a threshold on that same variable, so the sample structurally excludes its own extreme and the estimate is biased low with no bound. It is the neighbour of `instrument_output_read_as_subject_content` moved one step earlier: there a REPORT is complete for the reporter and read as complete for the consumer; here the truncation is in the ESTIMATOR'S OWN DOMAIN, and no amount of reading the instrument's output correctly recovers what the filter removed. SPECIMEN, both directions, 2026-09-01: the required floor's per-run cost inflation was estimated by pairing every claim that reported a cost in BOTH attempts of one tree -- median 1.053, p90 1.107. A claim that exceeds the cpu ceiling goes INTERRUPTED-BEFORE-VERDICT AND REPORTS NO COST, so the pairing dropped exactly the two most inflated observations; recovered from the ceiling itself, their inflation was at least 500/418 = 1.196, censored from above. THE SECOND HALF IS WHY THE ROW EXISTS: a peer lane, correcting a DIFFERENT defect in the same measurement, took the biased p90 by relay and derived an at-risk population of ONE -- an understatement arriving with the credibility of a retraction, which is the artifact nobody re-audits. Both counts were lower bounds and neither was labelled as one. RECOGNITION RULE: whenever an estimate is computed over rows that COMPLETED, RETURNED, PASSED, or were OBSERVED, ask what the incomplete rows would have contributed -- and if the incompleteness is caused by the very variable being estimated, the statistic is a floor and must be reported as one. The tell is a filter and an estimand naming the same quantity: cost estimated over rows that finished, latency over requests that did not time out, size over responses that were not truncated. REMEDY: report the censored bound rather than the sample statistic, and recover the excluded observations from the THRESHOLD they crossed, which is a real datum -- a row killed at 500ms is not missing, it is known to be above 500.)", evidence: [] } data restoration_promise_names_a_route_that_does_not_exist: RecurringFailureMode = RecurringFailureMode { identity: "restoration_promise_names_a_route_that_does_not_exist" as NonEmptyStr, authored: "**a restoration promise names a route that does not exist** (a mechanism withholds, defers or dormant-marks a population and tells the reader it becomes observable again under some named condition — and that condition names a RUN, LANE, CADENCE OR SWEEP the tree does not contain. Nothing refuses, because the promise is a string in a diagnostic rather than a citation anything resolves; the population is not silently dropped, which is what makes the class survive review — it is dropped WITH A RECEIPT, and the receipt is what stops anyone looking. **THE BOUNDARY AGAINST `unbacked_execution_claim` IS THE TENSE AND IT DECIDES THE REMEDY.** That class is prose asserting a relation that RUNS NOW; this is prose asserting a relation that WILL run — a future condition, so `git log --all -S` over the declaration form finds nothing to past-tense and the origin arms there do not apply. It is also not `absorbing_fallback`: nothing widens, the withhold is precise and correctly counted. It is 4b(3)'s trigger trap with the polarity inverted — there a trigger names LESS than the capability it restores and gets satisfied while the capability stays dead; here the trigger names a capability whose PRECONDITION IS ALREADY FALSE, so it can never be satisfied at all and the row waits forever in a state that reads as temporary. **SPECIMEN WITH A RECEIPT, measured 2026-09-01 on head 4c6c509e and unrepaired at authoring.** `v1_compiler.cli_run.required_floor_runner` `suppress_withheld` removes enrolled expected-red identities whose module sits outside the required gate, printing that their enrolment `becomes observable again when the gate roster admits the module or in the whole-corpus receipts run`. THAT RUN DOES NOT EXIST: the phrase occurs exactly once in the tree, inside the message that promises it, and the repository carries three workflows of which none is it. 39 identities across 23 modules sit under that promise. Each is enrolled on `v2.workflow.floor_expected_red`, whose own header states what an enrolment asserts — that the identity REACHES ITS SUBJECT AND ANSWERS, and that a row belongs there only while someone is fixing it — so every one of the 39 asserts `runs, fails, someone is fixing it` about a row no run reaches. The only surviving route by which one executes is the changed-witness override, i.e. somebody editing it. **THE POPULATION IS DERIVABLE AND IS DELIBERATELY NOT TRANSCRIBED HERE**: it is `v2.workflow.floor_expected_red` `floor_expected_red_roster` minus the identities whose module matches `v2.workflow.required_floor` `required_gate_prefixes` — two authorities and a set difference, so it re-derives instead of rotting. **THE SAME ROSTER'S HEADER ALREADY RECORDS THE ANCESTOR OF THIS MISTAKE**, which is why it is a class: 101 rows were held there as agreement while never reaching their subject, and were reclassified into `v2.workflow.floor_route_gap` on 2026-08-20 once `ExpectedRedArm` was taught to refuse `HostEffectRefused`, `HostToolUnresolved` and an interrupted budget. That repair closed the arm where a NON-VERDICT was read as agreement; this class is the same harm one step earlier, where a row never reaches an arm at all and a sentence promises it will. **RECOGNITION RULE: whenever a diagnostic says a withheld thing becomes observable again `in`/`under`/`by` some named run, grep the tree for that name and require an executing consumer — a workflow job, a scheduled entry point, an actuator argv. If the only occurrence is the promise itself, the population is dormant forever and the honest states are two: admit the row cannot be observed on any cadence, or delete the enrolment. **RUNG: 1 (mitigatable) — the withhold is counted and located, which is the whole of what holds. CEILING: 3, since `whether a named route exists` is decidable from the workflow and entry-point authorities the tree already carries. NEXT-RUNG TRIGGER, a CAPABILITY and not an artifact: restoration conditions expressed as a resolvable citation to an executing consumer rather than as prose, so a promise naming no route fails to compile — writing this particular sentence better retires nothing.)" as NonEmptyStr, evidence: [] } +data salience_instrument_blind_to_the_record_it_sizes: RecurringFailureMode = RecurringFailureMode { identity: "salience_instrument_blind_to_the_record_it_sizes" as NonEmptyStr, authored: "**a salience instrument blind to the record it sizes** (a change's SIZE IN THE DIFF is read to decide how much attention it deserves, and for one class of artifact that number is wrong by orders of magnitude, so a substantive edit is allocated a trivial edit's scrutiny. The instrument is not merely inaccurate, it is a SALIENCE instrument -- read FIRST, to decide where to look -- which is why the error does damage a plainly wrong number would not: nobody checks a figure they have already used to decide the thing is not worth checking. SPECIMEN, 2026-09-02, gunbc#10044: `git diff --stat -- dag/gunbc/rung_drop.dag` reported `1 insertion(+), 1 deletion(-)` for an edit adding 15,286 bytes, because every `RungDrop` in that corpus is ONE VERY LONG LINE. Three consecutive independent reviews described that hunk as a small wording tweak; a fourth, earlier one attributed it to `direct_call_arg_seam_v2_exemption`, A ROW THE DIFF DOES NOT TOUCH, because git's `@@` header names the declaration PRECEDING the hunk and the edited record is a single line. So one representation produced two distinct review failures -- under-sizing and misattribution -- and n=3 on the first within one pull request. THE APPROVALS WERE NOT WRONG ON WHAT THEY READ: each verdict is defensible over the prose rows the same diff carries, which render normally. What the count concealed is that the PR had three approvals and NO review coverage of the half that carried the receipt. A review tally is a claim about attention, and this instrument silently redirects attention before any reviewer forms a judgment. RUNG FOUND AT: silent wrongness, which is not a rung -- the misallocation leaves no trace, produces a green, and is indistinguishable from a reviewer who looked and found nothing. CEILING: mechanically preventable. The size of a record's diff can be made to track the size of its content, which is a representation question and decidable; it does not reach structural impossibility because nothing stops a future record from being authored as one long line again. NEXT-RUNG TRIGGER, AND IT IS DELIBERATELY NOT 'REVIEWERS SHOULD LOOK HARDER': a representation in which a record's diff size tracks its content size -- the long-line record broken across lines so a diff has hunks proportional to the change, or a review surface that reads the generated projection (`docs/design-ledgers.md` renders the same content as prose and diffs legibly) rather than the `.dag` line. 'Look harder' is advice, and advice is what a class gets when nobody wants to pay for the fix; it also cannot be discharged, since every reviewer meets the same instrument. RECOGNITION RULE: when a review's description of a hunk is smaller than the hunk, check whether the artifact's line structure and the change's content structure agree. Where one record is one line, EVERY size signal derived from lines -- the diffstat, the hunk header, a line-count budget, a review-effort heuristic -- is answering about the file's shape rather than the change's. The tell is a reviewer citing the enclosing declaration rather than the edited one.)" as NonEmptyStr, evidence: [] } + data non_verdict_disposition_surfaces_as_refusal: RecurringFailureMode = RecurringFailureMode { identity: "non_verdict_disposition_surfaces_as_refusal" as NonEmptyStr, authored: "**a non-verdict disposition surfaced as a refusal** (a gate COMPUTES the difference between THIS SUBJECT IS WRONG and I DID NOT FINISH LOOKING AT THIS SUBJECT, and then DISCARDS it at its aggregate boundary. THE NARROW CLAIM IS THE TRUE ONE AND AN EARLIER REVISION OF THIS ROW OVERSTATED IT: the floor does NOT conflate the two at the diagnostic grain, and saying so was falsifiable from the very log this row cites. It prints INTERRUPTED-BEFORE-VERDICT and COMPLETED-OVER-COST-REQUIREMENT as distinct typed diagnostics, and its summary carries `interrupted_before_verdict` and `completed_over_cost_requirement` as separate counters beside `failed`. A reader of the log can tell them apart perfectly. What cannot tell them apart is everything DOWNSTREAM of the fold -- the FloorRefused verdict, the red check, the dashboard cell, the merge gate -- each of which receives one bit and offers one affordance, which is to run it again. So the defect is not a missing distinction; it is a computed distinction erased on the way out, which is the worse shape of the two because the information exists and is thrown away. Its neighbour is `state_space_conflation` and its dual is `absorbing_fallback`: absorption widens an answer it could not compute, while this reports a computation it never completed AS an answer. The harm is specific and it is not the red itself: a refusal that names no wrong subject cannot be acted on, so the only available response is to run it again -- and a wall that is discharged by rerunning is not a wall. It teaches every consumer that its refusals are weather. SPECIMEN, 2026-09-02, gunbc.witness_floor_workflow required floor: the same commit, 9b00e24f592 on gunbc PR 9984, was run twice with no intervening edit -- ONE run id retried, Actions run 33604337589 attempts 1 and 2 -- and every counter below is addressed by ATTEMPT AND JOB rather than by the ordinals FIRST and RERUN, which name no instrument and do not distinguish this pair from run 33628404336, a different run on a later head. ATTEMPT 1, required-witnesses-floor job 100172868685, conclusion failure: planned=3486 executed=3486 passed=3404 failed=0 interrupted_before_verdict=4 completed_over_cost_requirement=3, verdict FloorRefused. ATTEMPT 2, required-witnesses-floor job 100189043027, conclusion success: planned=3486 executed=3486 passed=3411 failed=0 interrupted_before_verdict=0 completed_over_cost_requirement=0, verdict green. RE-DERIVE EACH SIDE FROM ITS OWN JOB with `gh api repos/gunb-ai/gunbc/actions/jobs//logs --allow-escape-sequences`, and NEVER with `gh run view --job --log`: measured on this very pair, the latter ANSWERS THE ATTEMPT-1 JOB ID WITH ATTEMPT 2's CONTENT -- its runner banner reads 09:03 where that job's own log begins 08:05, and it reports interrupted_before_verdict=0 completed_over_cost_requirement=0, which are attempt 2's counters. A reader trusting it would record 0 and 0 for BOTH attempts, see no disagreement, and destroy the specimen this row is built on. It is the sibling of `empty_capture_read_as_clean_result` arriving through a WRONG-CONTENT channel rather than an empty one, and it is the more dangerous direction because the answer is ~600 kB of well-formed, plausible log for the wrong subject rather than a zero-byte file anyone would question. Identical planned and executed populations, zero failures in BOTH runs, and the cost-arm rows simply absent the second time. THE SUBJECT OF THIS CLASS IS THE FOUR INTERRUPTED ROWS AND NOT ALL SEVEN, corrected after review: `completed_over_cost_requirement` names claims that REACHED A VERDICT and were then reclassified on cost -- the floor's own text says \"reached its verdict and then exceeded its budget ... cost=501ms EXACT ... This is a cost debt only -- it is not a defect\", and the rows carry outcome=completed_over_budget, which is to say they PASSED. Those three are a cost-debt fact about a gate that did judge, and folding them into a class about gates that did NOT judge would inflate the specimen by more than half while contradicting the model's own deliberate distinction. They are retained above only as the second half of the nondeterminism observation, where they belong: both arms of the cost machinery vary run to run on fixed bytes. THE SAME-HEAD PAIR IS THE POINT: a cross-head comparison would have required arguing that the intervening commit could not have touched cost accounting, and an argument about what a diff cannot do is exactly what gets overturned; holding the bytes fixed by construction leaves nothing to argue. The disposition at issue is named non-verdict by the floor itself: INTERRUPTED-BEFORE-VERDICT says the deadline preempted the witness so whether it passes is UNKNOWN, and the row's cost is reported as above the ceiling with NO UPPER BOUND -- a bound, not a near miss. So on those four rows the gate is not reversing a judgment between runs; it is failing to REACH one and rendering that as a refusal. That is the whole class: FOUR rows in this specimen, and the arithmetic of the refusal is that four undecided rows were sufficient to refuse a run in which zero claims failed. THE COST HALF OF THIS STORY IS NOT THIS ROW'S AND IS NOT RE-DERIVED HERE. That a claim's measured cpu-ms is unstable on a shared runner -- and therefore that `completed_over_cost_requirement` varies run to run on fixed bytes -- is the declared standing drop `gunbc.rung_drop` `floor_cost_contention_verdict`, whose restoration trigger is a claim-owned cost basis invariant or bounded by construction across the admitted envelopes. That row also settles what happens at the boundary: both terminal arms STAY required reds and are distinct. This row takes no position on whether a cost verdict qualifies a claim; it says only that the aggregate discards which arm fired. Filing the cost instability separately would be a second authority over one fact. RUNG FOUND AT: mitigatable -- the dispositions are typed, located and counted at the fold, and the floor's own diagnostic text states plainly that the verdict is unknown rather than claiming a failure. What it does not do is carry that distinction past its own summary line. CEILING: structurally guaranteed. Whether a claim reached a verdict is decidable and already decided -- the floor holds the disposition -- so this is DESIGN section 5's wall after grounding, not a decidability limit. NEXT-RUNG TRIGGER, two conjuncts because either alone leaves the class where it is: (i) a required-gate verdict in which non-verdict dispositions are a THIRD outcome distinct from pass and fail -- DISTINCT IN DIAGNOSIS, NEVER IN WHETHER THE LINE STOPS. The third outcome must still block the gate: a run that did not finish looking has established nothing, and letting it through would be a widening arm of exactly the kind DESIGN section 5 forbids, trading a refusal that names no subject for no refusal at all. What must change is what the refusal SAYS -- an undecided row is reported as undecided and is discharged by making the claim reach a verdict, never by a rerun that happens to land under the ceiling. The defect is the conflation, not the stopping; AND (ii) a cost admission whose budget is a property of the CLAIM rather than of the run's host speed, sufficient that the same bytes yield the same disposition across runs -- without it, arm (i) merely relabels an intermittent outcome and the population that must be resolved keeps changing underneath its own roster. RECOGNITION RULE: ask what the gate reports when it could not measure, and compare it with what it reports when it measured and found the subject wrong. If those two are the same value, the gate's refusals do not name a subject, and the tell is behavioural: people begin rerunning it rather than reading it.)" as NonEmptyStr, evidence: [] } data empty_capture_read_as_clean_result: RecurringFailureMode = RecurringFailureMode { identity: "empty_capture_read_as_clean_result" as NonEmptyStr, authored: "**an empty capture read as a clean result** (an instrument refuses on one stream while the reader keeps the other, so the capture is EMPTY and the empty capture is consumed as a finding of nothing. It is a specialization of `blank is not evidence` narrowed to a mechanism that recurs: the refusal is well-formed and honest, it simply travels on a channel nobody kept. SPECIMEN, 2026-09-02, reading a failing CI job during gunbc PR 9984: `gh run view --job --log > f` writes ZERO BYTES and puts `run is still in progress; logs will be available when it is complete` on stderr, so the redirect captures nothing while the job is already recorded as completed failure. A grep for `panicked` or `FAILED` over that file then reports no failures for a job that failed. Two adjacent instances in the same call: `gh api .../actions/jobs//logs` writes zero bytes without --allow-escape-sequences, again refusing on stderr; and the same subcommand takes a JOB id rather than a run id, where the wrong id also yields nothing useful. THE ASYMMETRY IS WHAT MAKES IT RECUR: the failure direction is always benign. An empty capture never manufactures a false alarm, only a false all-clear, so nothing in the reader's experience trains them to check it. RUNG FOUND AT: mitigatable at the reading site -- measuring the capture's size before consuming it converts the silent case into a loud one -- and silent wrongness wherever that check is absent, which is the default. CEILING: mechanically preventable for instruments this repository owns, since a reader that must consume a typed outcome cannot consume a byte count of zero as a verdict; it stays mitigatable for third-party command-line tools, whose stream discipline is outside the modeled guarantee and is observed rather than proven. NEXT-RUNG TRIGGER: a repository instrument-reading surface whose result type distinguishes READ SUCCESSFULLY AND FOUND NOTHING from DID NOT READ, sufficient that a zero-byte capture cannot inhabit the first -- at which point the diligence rule below is subsumed for every call routed through it, and remains diligence for calls that are not. RECOGNITION RULE: before believing a quiet result, ask which stream the instrument refuses on and which stream you kept. If a refusal and a clean result produce the same bytes in the channel you read, you have no observation. The mechanical form, and it is a two-step rather than a byte-count rule, because inferring failure from size alone fabricates a refusal for the query that legitimately returned nothing: a zero-byte capture is NOT A VERDICT IN EITHER DIRECTION, so consult the instrument's typed status -- its exit code, and the stream its refusal travels on -- before consuming the emptiness. Where the status says the instrument ran and the capture is empty, that is FOUND NOTHING and is a real observation; where the status says it refused, or the status is unavailable, the capture is DID NOT READ and no conclusion may be drawn from it. The failure of the original habit is not that it read zero as one of those two, it is that it read zero WITHOUT ASKING which.)" as NonEmptyStr, evidence: [] } @@ -296,6 +298,7 @@ data recurring_failure_mode_roster: List = [ liveness_probe_read_as_currency, non_verdict_disposition_surfaces_as_refusal, empty_capture_read_as_clean_result, + salience_instrument_blind_to_the_record_it_sizes, ambient_process_state_read_by_a_concurrent_reader, predicate_vacuously_true_on_an_empty_domain, check_subject_narrower_than_its_declared_claim, diff --git a/docs/design-ledgers.md b/docs/design-ledgers.md index 1cb5e65086e..63503f8e523 100644 --- a/docs/design-ledgers.md +++ b/docs/design-ledgers.md @@ -65,6 +65,7 @@ The landing measurement partitions the 31 parser-visible identities into **2 cit - **a liveness probe read as currency** (a probe answers whether a thing is RUNNING and its answer is consumed as whether it is running the CURRENT definition. The two questions differ only for a process that outlived the definition it was started from, which is the one state a converge is supposed to repair, so the conflation is invisible on every host that has never drifted and load-bearing on exactly the hosts that have. Its neighbour is `state_space_conflation`, and what separates it is that NOTHING HERE IS OVERLOADED: each probe answers its own subject correctly and the loss happens in the JOIN, where a reader picks the observation nearest the question instead of the one that answers it. SPECIMEN, 2026-09-02, gunbc.spark.serving_membership: serving convergence carries an activity member observed through `systemctl --user is-active`, whose four states are running, stopped, no-such-unit and unobserved, and a registration member whose digest is the unit content ON DISK. A host whose unit file was rewritten and whose manager was reloaded converges at BOTH addresses -- the disk holds the desired bytes, the manager has consumed them, the unit is active -- while the process still serves the definition it was started from. No member disagrees, so no remedy is planned, and the drift is reported as convergence. THE HARM IS NOT THE STALE PROCESS, IT IS THAT NOTHING CAN EVER SELECT ITS REMEDY: a restart has no observation to be conditioned on, which is why gunbc.spark.serving_realization carried its only restart as an unconditional third command welded into EnableSystemUnit, restarting every host on every converge because it could not tell which host needed it. RUNG FOUND AT: silent wrongness, which is not a rung -- a stale server and a converged one produced the same verdict with no diagnostic. THE STATE PREDATES THE CHANGE THAT FILED THIS ROW, stated because an honest declaration of a pre-existing gap is indistinguishable, to a reader scanning a diff for regressions, from a confession of a new one -- and it has now been read that way twice. The convergence path acquired this class when activity was modeled on `systemctl --user is-active` in gunbc.spark.serving_observed_members, whose four states name whether a unit runs and never which definition it runs. The PR that split the fused restart neither introduced this state nor closed it; it named it. IT IS STILL THERE ON THE CONVERGENCE PATH, and an earlier revision of this row said MITIGATABLE, which was rung inflation caught in review: `spark_serving_reactivation_remedy` refuses with a typed located cause, but it has NO PRODUCTION CALLER, so nothing routes a convergence through it and the operational path still reports convergence without the refusal. A rung is the minimum across a class's in-scope paths, and a selector nothing calls moves none of them. WHAT THIS CHANGE DID ESTABLISH is the vocabulary and a selector whose refusal is executed against FIXTURE inputs -- a declared boundary, not a rung, and named as one so it cannot be cited as coverage. THE ABSENCE OF THE CALLER IS A DECISION, NOT AN OMISSION: a production caller is a path that can emit a restart, and until an admission fence, a drain protocol and a lease population exist, a restart emitted on a live host is the destructive convergence the gate exists to prevent -- so wiring it early would trade an honest gap for an unguarded one. CEILING: structurally guaranteed. The question is decidable -- a running invocation either carries the definition identity it was started from or does not -- so this is DESIGN section 5's wall AFTER GROUNDING, waiting on its single authority rather than on a decidability argument. NEXT-RUNG TRIGGER, AND IT IS TWO THINGS RATHER THAN ONE, because either alone leaves the class exactly where it is: (i) an observation that joins a unit's RUNNING INVOCATION to the definition it was started from, sufficient to mint `RunningDefinitionObserved` for a live Spark -- nothing smaller counts, since a probe that reads the file, the manager's loaded definition, or liveness answers a different subject and would leave the capability dead while looking like its trigger; AND (ii) a guarded producer that routes `ReactivationUndecidable` into the convergence verdict, so a host whose running definition cannot be established stops instead of reporting converged. Naming only (i) would be the grain mismatch section 4b(3) warns about: the trigger would be satisfied by a probe landing while the convergence path stayed silent, which is the whole defect. THE OBSERVED-VERSUS-DESIRED DRIFT ON BOTH SPARKS IS A CANDIDATE CONSEQUENCE OF THIS CLASS AND NOT A FINDING: nobody has read a host back, and candidate is the correct word until someone does. RECOGNITION RULE: when a converge concludes from an observation, ask what the observation would say about a subject that is CORRECT BUT STALE -- if the probe answers the same for a current and an outdated instance, it is a liveness reading standing where a currency reading was needed. The tell is a remedy with no condition: an operation that is issued every time, or never, because no observation can distinguish the state it repairs.) - **a non-verdict disposition surfaced as a refusal** (a gate COMPUTES the difference between THIS SUBJECT IS WRONG and I DID NOT FINISH LOOKING AT THIS SUBJECT, and then DISCARDS it at its aggregate boundary. THE NARROW CLAIM IS THE TRUE ONE AND AN EARLIER REVISION OF THIS ROW OVERSTATED IT: the floor does NOT conflate the two at the diagnostic grain, and saying so was falsifiable from the very log this row cites. It prints INTERRUPTED-BEFORE-VERDICT and COMPLETED-OVER-COST-REQUIREMENT as distinct typed diagnostics, and its summary carries `interrupted_before_verdict` and `completed_over_cost_requirement` as separate counters beside `failed`. A reader of the log can tell them apart perfectly. What cannot tell them apart is everything DOWNSTREAM of the fold -- the FloorRefused verdict, the red check, the dashboard cell, the merge gate -- each of which receives one bit and offers one affordance, which is to run it again. So the defect is not a missing distinction; it is a computed distinction erased on the way out, which is the worse shape of the two because the information exists and is thrown away. Its neighbour is `state_space_conflation` and its dual is `absorbing_fallback`: absorption widens an answer it could not compute, while this reports a computation it never completed AS an answer. The harm is specific and it is not the red itself: a refusal that names no wrong subject cannot be acted on, so the only available response is to run it again -- and a wall that is discharged by rerunning is not a wall. It teaches every consumer that its refusals are weather. SPECIMEN, 2026-09-02, gunbc.witness_floor_workflow required floor: the same commit, 9b00e24f592 on gunbc PR 9984, was run twice with no intervening edit -- ONE run id retried, Actions run 33604337589 attempts 1 and 2 -- and every counter below is addressed by ATTEMPT AND JOB rather than by the ordinals FIRST and RERUN, which name no instrument and do not distinguish this pair from run 33628404336, a different run on a later head. ATTEMPT 1, required-witnesses-floor job 100172868685, conclusion failure: planned=3486 executed=3486 passed=3404 failed=0 interrupted_before_verdict=4 completed_over_cost_requirement=3, verdict FloorRefused. ATTEMPT 2, required-witnesses-floor job 100189043027, conclusion success: planned=3486 executed=3486 passed=3411 failed=0 interrupted_before_verdict=0 completed_over_cost_requirement=0, verdict green. RE-DERIVE EACH SIDE FROM ITS OWN JOB with `gh api repos/gunb-ai/gunbc/actions/jobs//logs --allow-escape-sequences`, and NEVER with `gh run view --job --log`: measured on this very pair, the latter ANSWERS THE ATTEMPT-1 JOB ID WITH ATTEMPT 2's CONTENT -- its runner banner reads 09:03 where that job's own log begins 08:05, and it reports interrupted_before_verdict=0 completed_over_cost_requirement=0, which are attempt 2's counters. A reader trusting it would record 0 and 0 for BOTH attempts, see no disagreement, and destroy the specimen this row is built on. It is the sibling of `empty_capture_read_as_clean_result` arriving through a WRONG-CONTENT channel rather than an empty one, and it is the more dangerous direction because the answer is ~600 kB of well-formed, plausible log for the wrong subject rather than a zero-byte file anyone would question. Identical planned and executed populations, zero failures in BOTH runs, and the cost-arm rows simply absent the second time. THE SUBJECT OF THIS CLASS IS THE FOUR INTERRUPTED ROWS AND NOT ALL SEVEN, corrected after review: `completed_over_cost_requirement` names claims that REACHED A VERDICT and were then reclassified on cost -- the floor's own text says "reached its verdict and then exceeded its budget ... cost=501ms EXACT ... This is a cost debt only -- it is not a defect", and the rows carry outcome=completed_over_budget, which is to say they PASSED. Those three are a cost-debt fact about a gate that did judge, and folding them into a class about gates that did NOT judge would inflate the specimen by more than half while contradicting the model's own deliberate distinction. They are retained above only as the second half of the nondeterminism observation, where they belong: both arms of the cost machinery vary run to run on fixed bytes. THE SAME-HEAD PAIR IS THE POINT: a cross-head comparison would have required arguing that the intervening commit could not have touched cost accounting, and an argument about what a diff cannot do is exactly what gets overturned; holding the bytes fixed by construction leaves nothing to argue. The disposition at issue is named non-verdict by the floor itself: INTERRUPTED-BEFORE-VERDICT says the deadline preempted the witness so whether it passes is UNKNOWN, and the row's cost is reported as above the ceiling with NO UPPER BOUND -- a bound, not a near miss. So on those four rows the gate is not reversing a judgment between runs; it is failing to REACH one and rendering that as a refusal. That is the whole class: FOUR rows in this specimen, and the arithmetic of the refusal is that four undecided rows were sufficient to refuse a run in which zero claims failed. THE COST HALF OF THIS STORY IS NOT THIS ROW'S AND IS NOT RE-DERIVED HERE. That a claim's measured cpu-ms is unstable on a shared runner -- and therefore that `completed_over_cost_requirement` varies run to run on fixed bytes -- is the declared standing drop `gunbc.rung_drop` `floor_cost_contention_verdict`, whose restoration trigger is a claim-owned cost basis invariant or bounded by construction across the admitted envelopes. That row also settles what happens at the boundary: both terminal arms STAY required reds and are distinct. This row takes no position on whether a cost verdict qualifies a claim; it says only that the aggregate discards which arm fired. Filing the cost instability separately would be a second authority over one fact. RUNG FOUND AT: mitigatable -- the dispositions are typed, located and counted at the fold, and the floor's own diagnostic text states plainly that the verdict is unknown rather than claiming a failure. What it does not do is carry that distinction past its own summary line. CEILING: structurally guaranteed. Whether a claim reached a verdict is decidable and already decided -- the floor holds the disposition -- so this is DESIGN section 5's wall after grounding, not a decidability limit. NEXT-RUNG TRIGGER, two conjuncts because either alone leaves the class where it is: (i) a required-gate verdict in which non-verdict dispositions are a THIRD outcome distinct from pass and fail -- DISTINCT IN DIAGNOSIS, NEVER IN WHETHER THE LINE STOPS. The third outcome must still block the gate: a run that did not finish looking has established nothing, and letting it through would be a widening arm of exactly the kind DESIGN section 5 forbids, trading a refusal that names no subject for no refusal at all. What must change is what the refusal SAYS -- an undecided row is reported as undecided and is discharged by making the claim reach a verdict, never by a rerun that happens to land under the ceiling. The defect is the conflation, not the stopping; AND (ii) a cost admission whose budget is a property of the CLAIM rather than of the run's host speed, sufficient that the same bytes yield the same disposition across runs -- without it, arm (i) merely relabels an intermittent outcome and the population that must be resolved keeps changing underneath its own roster. RECOGNITION RULE: ask what the gate reports when it could not measure, and compare it with what it reports when it measured and found the subject wrong. If those two are the same value, the gate's refusals do not name a subject, and the tell is behavioural: people begin rerunning it rather than reading it.) - **an empty capture read as a clean result** (an instrument refuses on one stream while the reader keeps the other, so the capture is EMPTY and the empty capture is consumed as a finding of nothing. It is a specialization of `blank is not evidence` narrowed to a mechanism that recurs: the refusal is well-formed and honest, it simply travels on a channel nobody kept. SPECIMEN, 2026-09-02, reading a failing CI job during gunbc PR 9984: `gh run view --job --log > f` writes ZERO BYTES and puts `run is still in progress; logs will be available when it is complete` on stderr, so the redirect captures nothing while the job is already recorded as completed failure. A grep for `panicked` or `FAILED` over that file then reports no failures for a job that failed. Two adjacent instances in the same call: `gh api .../actions/jobs//logs` writes zero bytes without --allow-escape-sequences, again refusing on stderr; and the same subcommand takes a JOB id rather than a run id, where the wrong id also yields nothing useful. THE ASYMMETRY IS WHAT MAKES IT RECUR: the failure direction is always benign. An empty capture never manufactures a false alarm, only a false all-clear, so nothing in the reader's experience trains them to check it. RUNG FOUND AT: mitigatable at the reading site -- measuring the capture's size before consuming it converts the silent case into a loud one -- and silent wrongness wherever that check is absent, which is the default. CEILING: mechanically preventable for instruments this repository owns, since a reader that must consume a typed outcome cannot consume a byte count of zero as a verdict; it stays mitigatable for third-party command-line tools, whose stream discipline is outside the modeled guarantee and is observed rather than proven. NEXT-RUNG TRIGGER: a repository instrument-reading surface whose result type distinguishes READ SUCCESSFULLY AND FOUND NOTHING from DID NOT READ, sufficient that a zero-byte capture cannot inhabit the first -- at which point the diligence rule below is subsumed for every call routed through it, and remains diligence for calls that are not. RECOGNITION RULE: before believing a quiet result, ask which stream the instrument refuses on and which stream you kept. If a refusal and a clean result produce the same bytes in the channel you read, you have no observation. The mechanical form, and it is a two-step rather than a byte-count rule, because inferring failure from size alone fabricates a refusal for the query that legitimately returned nothing: a zero-byte capture is NOT A VERDICT IN EITHER DIRECTION, so consult the instrument's typed status -- its exit code, and the stream its refusal travels on -- before consuming the emptiness. Where the status says the instrument ran and the capture is empty, that is FOUND NOTHING and is a real observation; where the status says it refused, or the status is unavailable, the capture is DID NOT READ and no conclusion may be drawn from it. The failure of the original habit is not that it read zero as one of those two, it is that it read zero WITHOUT ASKING which.) +- **a salience instrument blind to the record it sizes** (a change's SIZE IN THE DIFF is read to decide how much attention it deserves, and for one class of artifact that number is wrong by orders of magnitude, so a substantive edit is allocated a trivial edit's scrutiny. The instrument is not merely inaccurate, it is a SALIENCE instrument -- read FIRST, to decide where to look -- which is why the error does damage a plainly wrong number would not: nobody checks a figure they have already used to decide the thing is not worth checking. SPECIMEN, 2026-09-02, gunbc#10044: `git diff --stat -- dag/gunbc/rung_drop.dag` reported `1 insertion(+), 1 deletion(-)` for an edit adding 15,286 bytes, because every `RungDrop` in that corpus is ONE VERY LONG LINE. Three consecutive independent reviews described that hunk as a small wording tweak; a fourth, earlier one attributed it to `direct_call_arg_seam_v2_exemption`, A ROW THE DIFF DOES NOT TOUCH, because git's `@@` header names the declaration PRECEDING the hunk and the edited record is a single line. So one representation produced two distinct review failures -- under-sizing and misattribution -- and n=3 on the first within one pull request. THE APPROVALS WERE NOT WRONG ON WHAT THEY READ: each verdict is defensible over the prose rows the same diff carries, which render normally. What the count concealed is that the PR had three approvals and NO review coverage of the half that carried the receipt. A review tally is a claim about attention, and this instrument silently redirects attention before any reviewer forms a judgment. RUNG FOUND AT: silent wrongness, which is not a rung -- the misallocation leaves no trace, produces a green, and is indistinguishable from a reviewer who looked and found nothing. CEILING: mechanically preventable. The size of a record's diff can be made to track the size of its content, which is a representation question and decidable; it does not reach structural impossibility because nothing stops a future record from being authored as one long line again. NEXT-RUNG TRIGGER, AND IT IS DELIBERATELY NOT 'REVIEWERS SHOULD LOOK HARDER': a representation in which a record's diff size tracks its content size -- the long-line record broken across lines so a diff has hunks proportional to the change, or a review surface that reads the generated projection (`docs/design-ledgers.md` renders the same content as prose and diffs legibly) rather than the `.dag` line. 'Look harder' is advice, and advice is what a class gets when nobody wants to pay for the fix; it also cannot be discharged, since every reviewer meets the same instrument. RECOGNITION RULE: when a review's description of a hunk is smaller than the hunk, check whether the artifact's line structure and the change's content structure agree. Where one record is one line, EVERY size signal derived from lines -- the diffstat, the hunk header, a line-count budget, a review-effort heuristic -- is answering about the file's shape rather than the change's. The tell is a reviewer citing the enclosing declaration rather than the edited one.) - **ambient process state read by a concurrent reader** (a fact that is really a PARAMETER is instead read from a mutable cell the whole process shares -- the working directory, an env var, a global -- while the runtime schedules concurrent readers of it, so the answer a reader gets depends on which writer last ran and the wrong answer is SILENT: the read succeeds, returns a well-formed value, and is simply about a different subject than the reader meant. INVALID STATE: any reader resolving a relative name against ambient state that a concurrent writer may move mid-read. WHY IT SURVIVES REVIEW: each site is locally correct -- it sets the state to the right value before its own work and restores it after -- and the defect exists only in the PRODUCT of sites, which no single diff shows. SPECIMEN (PR #10024, v1_compiler.cli_run): 38 forward set_current_dir sites in one test binary that cargo runs multi-threaded, with the idiom prior = current_dir(); chdir(ws); work; chdir(prior) -- so one test's RESTORE hands the cwd back while a sibling is mid-read. The victim population is not the writers, it is every non-ignored test in the binary that resolves a relative path during the window; the file's own entry_admission_tests note already records the measurement, a DIFFERENT pair of tests failing on each run of identical code (4/1 then 3/2). Seven of the sites were reachable from the required unit lane and in all seven the chdir was DEAD -- every consumer already took an explicit root -- i.e. residue from an unfinished parameterisation, which is the expected shape once the migration is half done. THE CENSUS THAT MATTERS IS REACHABILITY, NOT THE LITERAL AND THE RANGE THAT MATTERS IS THE POPULATION'S, NOT THE FILE'S: the first cut of that gate scanned the one file its author was editing while declaring the whole lane as its subject, which is a selection view read as a population and puts the reported rung above the executed evidence (DESIGN section 4b(1) reads a class's rung as the MINIMUM across in-scope paths). Widening it to every source file of the lane's crate cost a change of input plus a linear-time call extractor -- the per-name matcher is quadratic in lines times declarations and does not survive the wider corpus -- and it surfaced a second trap immediately: resolving cross-file call edges by BARE NAME merges every declaration of a shared spelling, and this crate declares main 18 times, so one production seed bridged into hundreds of unrelated tests. That arm is the absorbing fallback again, in the guard built to catch it. What survives is intra-file edges plus cross-file edges only for names declared exactly once, with the ambiguous remainder pinned by IDENTITY rather than assumed empty. Six forms of the same defect were found in this one guard; they are rostered as their own class at check_subject_narrower_than_its_declared_claim, which this row is a specimen of rather than an authority for: all seven reached the mutation through a helper, so a body-local grep for the call is green over every instance of the class and stays green the moment a writer is moved one call deeper -- the guard must close transitively over call edges. A SECOND, DISTINCT ARM, kept apart because it has a different remedy: memoizing a value DERIVED from the ambient state (workspace_root's OnceLock over a .git-ancestor walk from cwd) freezes whichever writer won the race for the whole process. On this specimen that arm's population is EMPTY -- every forward site targets the checkout root and 37 of 38 force the memo from the pristine cwd one line above their own call -- so it is reported UNSUPPORTED rather than disproven (see a_structural_possibility_is_not_an_occurring_phenomenon); ONE chdir to a non-checkout path arms it, and a fixture that runs git init is such a path. RUNG: found BELOW the ladder, since a redirected read is a silently wrong answer with no typed refusal anywhere. CEILING 4, structurally impossible, and decidable: state passed as a parameter has no shared cell to race on, so the invalid state has no constructor. The intermediate rung is a reachability gate over the guarded population (no_gating_test_reaches_a_process_cwd_mutator), which keeps the state ABSENT but not unwritable; its next-rung trigger is the capability that every reader in the population resolves from an explicitly passed root, after which the ambient write has no consumer and the call can be removed rather than merely counted. REVIEW TELL: a diff that sets shared process state and restores it, in code the runtime may run concurrently -- the restore is the tell, because it is what makes the site look self-contained. - **a predicate vacuously true on an empty domain** (a universally-quantified check -- `all`, `every`, `iter().all(..)`, a for-loop that only ever narrows a flag, a SQL NOT EXISTS -- is satisfied by the IDENTITY ELEMENT when its domain is empty, so the check answers TRUE about a population it never examined. INVALID STATE: a guard whose verdict is read as evidence while its domain is empty, with nothing establishing that the domain was non-empty. WHY IT SURVIVES REVIEW, and this is what separates it from an ordinary off-by-one: the code READS as a check, it EXECUTES, it is REACHED, and on every input a reviewer imagines the domain is non-empty and the predicate genuinely discriminates. The vacuous case is not a value in the data -- IT IS PRODUCED BY A BOUNDARY, so it appears only at the end of a line, the end of a buffer, a remainder shorter than the window, an empty selection, the last chunk. Those are exactly the inputs a hand-written fixture omits, and exactly the inputs a real corpus eventually contains. SPECIMEN (PR #10033, measured): a raw-string scanner tested its closing delimiter with `chars[i+1..].iter().take(hashes).all(|h| *h == '#')`. Over a remainder shorter than `hashes` that iterator yields fewer items and `all` is vacuously TRUE, so a quote near end-of-line closes a hash-delimited raw string that is still open, and the rest of the fixture is handed back to the code scanner -- reintroducing the very brace-depth skew the raw-string handling existed to prevent, silently and in the UNDER-approximating direction. The repair is a length bound, not a different predicate: `i + 1 + hashes <= len` conjoined with the same `all`, which leaves the zero-hash case (a raw string with no hashes) trivially satisfied as it must be. RECOGNITION RULE, and it generalises past scanners: A PREDICATE THAT IS TRUE ON THE EMPTY CASE IS NOT A CHECK UNTIL ITS DOMAIN IS PROVEN NON-EMPTY. Ask of every universal: what does this return when the collection is empty, and can the boundary produce empty? If the answer is TRUE and YES, the domain bound is part of the predicate and its absence is the defect. BOUNDED AGAINST THREE NEIGHBOURS, none of whose recognition rules find it. `executed_conjunct_discriminates_nothing` is about a CORPUS that never constructs the falsifying population, so its remedy is a fixture; here the falsifying population is unconstructable at the boundary by the shape of the quantifier itself, and the remedy is a bound on the domain. `empty_observation_narrow` is an OBSERVATION that could not express what changed being rendered as a verdict -- one layer up, about evidence rather than about a predicate's truth value. `incidental_denominator_as_wall` is a guard that holds by a coincidence upstream that genuinely holds; here nothing holds at all, the guard simply returns TRUE. RUNG: found OUTSIDE the ladder, because the wrong answer is silent and well-formed -- the check reports satisfaction. CEILING 4 and decidable per-site: a windowed comparison whose window length is carried in its type cannot be asked about a short remainder. The intermediate rung is an executed control that exercises the BOUNDARY case specifically, asserted on a synthetic input rather than on the live corpus, so it keeps discriminating power when the corpus changes. - **a check's subject is narrower than the claim it is cited for** (EVERY DEGREE OF FREEDOM IN A CHECK'S SUBJECT -- which call sites, which spellings, which lexical states, which domain, which files -- is a place its coverage can silently be narrower than the population it is read as covering. INVALID STATE: an executing, reached, green check whose declared subject is a strict superset of the population its implementation actually ranges over. RUNG: outside the ladder, because the gap is invisible from the check's own result -- a wall whose range is too narrow is green for the same reason a wall over a clean corpus is green, and DESIGN section 4b(1) reads a class's rung as the MINIMUM across its in-scope paths, so a check like this INFLATES the rung of everything it is cited for. THE ROUTE IS WHAT MAKES IT INVISIBLE: the narrowing is never a decision, it is the shape of the first implementation -- an author edits one file, so the scanner reads one file; the corpus spells the call one way, so the seed matches one spelling; the fixture has no raw strings, so the projection has no raw-string state; the window is never short, so the universal is never vacuous. Each is locally reasonable and none is written down, which is why the gap survives the review that reads the check's LOGIC and never asks what its INPUT was. SPECIMEN, and the reason this is a class rather than a bug: ONE guard in PR #10033 carried SIX instances -- (1) a body-local grep for the mutator, defeated by moving the call one helper deeper; (2) a seed matching only the fully-qualified path, defeated by `use std::env;`; (3) a lexical projection with no raw-string state, so an unbalanced brace in a fixture silently shifted every containment decision after it; (4) `take(n).all(..)`, vacuously true over a short remainder (rostered as predicate_vacuously_true_on_an_empty_domain); (5) a scan ranging over ONE FILE while the scope declaration claimed every test the required lane runs -- found by review, not by the author; and (6) after widening, cross-file call edges resolved by BARE NAME, which merges every declaration of a shared spelling (18 `main`s in that crate) and bridged one production seed into hundreds of unrelated tests. Four were found by the author, one by review, one by the widening itself, and the sixth is the mirror of the fifth: over-broad and under-broad are the same defect measured in opposite directions, and both are reported as coverage. RECOGNITION RULE: name the check's SUBJECT and the check's RANGE as two separate sentences, then ask what is in the first and not the second. If the answer is anything, the claim is the narrower one until the range is widened. The review-side tell is a scope comment that quantifies (`every test`, `all modules`, `any call site`) beside an implementation that names ONE input. BOUNDED AGAINST TWO NEIGHBOURS. selection_view_read_as_population is a set produced by a MEASUREMENT threshold and read as the population, arriving as the fix for a real objection; here nothing is measured and nothing is filtered -- the range was simply never as wide as the sentence beside it. incidental_denominator_as_wall is a guard that holds by an upstream coincidence that genuinely holds; here the guard holds over a population it never read. CEILING 3, structurally guaranteed, and decidable: a check whose subject is a VALUE it receives rather than a scope its implementation re-derives cannot range over less than it claims, because the claim and the range are then one object. NEXT-RUNG TRIGGER, named as a capability: the guarded population is passed to the check as a constructed value carrying its own extent, sufficient that no check can name a subject its input does not contain. WHAT IS NOT CLAIMED: that widening always wins. Narrowing the DECLARATION to match the range is equally honest and is sometimes correct -- what is forbidden is leaving the two different, and the choice between them is a cost measurement, not a preference. On this specimen widening cost one change of input plus one linear-time rewrite of a quadratic edge builder, and the wider check ran FASTER than the narrow one, which is what settled it.)