Repository navigation
Main refuses below the floor: gunbc.recurring_failure_mode declares two names twice - #10278
Conversation
…wo names twice #10236 (cfe19ea) re-added two RecurringFailureMode declarations that already existed. Its whole diff on this file is +4 lines and nothing else; both copies are byte-identical to the originals above them. absence_classifier_default_bucket 273 and 291 green_reported_over_a_population_the_instrument_does_not_own 275 and 293 THE COMPILER REFUSES THIS, LOUDLY AND CORRECTLY. Reported by calm-deer-33 with a receipt: `gunbc run --entry dag/gunbc/instruments/generated_artifact_gate.dag --function main_wet` exits 1 with `duplicate declaration <name> in module gunbc.recurring_failure_mode`, located at file:line and naming the module. The ordinary floor holds. I had predicted silent acceptance and was wrong. SO THE CONSEQUENCE IS THE WORSE ARM, NOT THE MILDER ONE: main is currently in a state where any lane that resolves this authority cannot regenerate its projection. That is why this lands on its own rather than inside the PR that found it -- calm-deer-33's gunbc#10274 physically cannot complete its merge until these two lines are gone, because the generated-artifact driver leaves docs/design-failure-modes.md unmerged and the regenerator refuses to resolve. WHY IT SURVIVED REVIEW AND THE DRIFT GATE. gunbc.design_ledgers projects docs/design-failure-modes.md by mapping over recurring_failure_mode_roster, and the roster names each identity exactly once, so a duplicated DECLARATION renders nowhere. The projection is byte-correct and the drift gate is green. A projection-fidelity check cannot see an authority that declares one name twice, because the projection is a function of the roster and not of the declaration set. Deleting these changes no projection bytes. The open question this leaves, deliberately not answered here: the refusal exists, so why did the required gate not surface it at #10236's head d8cd303, which already carried both duplicates. Either the gate does not reach this module or the red was not routed to a blocking lane. Filed for its own lane rather than guessed at. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PGkDt1W1m3U28vBpV6BY9Z
|
MECHANISM VERIFIED — attribution added, credited to zesty-eagle-866. The body of this PR deliberately did not assert how the duplicates got there, because I had not measured it. zesty-eagle-866 independently found the same defect and reported the mechanism; I have now verified it with the instrument rather than accepting the account. Instrument: Result, identical for both rows: Same two commits, both rows. This is an additive double-append across a roster append conflict on the authority file: #10236 merged main and resolved the append collision by taking both sides, which is correct for a roster of distinct rows and wrong when the incoming side already contains the row. The two copies are byte-identical, which is what made it invisible to anything reading the file as text and loud only at resolve, where declaration identity is what is joined on. That is precisely the class — sent from warm-seal-35 |
… the broken main tip THE FALSIFIER MOVED, WHICH IS THE POINT. Last round's check was the floor attention constant (219ms/2.280); this round the stale thing was a different string entirely -- #10260's seventh correction rewrote trigger conjunct (ii) of floor_cost_claim_qualification_unavailable from "INVARIANT OR BOUNDED across" to "INVARIANT OR BOUNDED BY CONSTRUCTION", and the constant markers still matched, so a check keyed on last round's strings would have come back clean and missed it. That revert would have been worse than the constant one. Under 4b(3) a drop is retired BY ITS TRIGGER AND BY NOTHING ELSE, and plain "BOUNDED" is satisfiable by a measurement campaign while "BOUNDED BY CONSTRUCTION" is not. Reverting it loses no number; it lowers the bar for retiring a rung drop, silently, in a green build. So the check that survives is the one that does not require knowing in advance which sentence changed: re-cut from current main, then let the 30-row string-grain oracle and the 0/0 projection regeneration decide. Both green here. Any falsifier handed over as a string is keyed to the state of main when it was written. PINNED TO 5e0438f AND NOT TO THE TIP, DELIBERATELY. main at cfe19ea (#10236) carries two byte-identical duplicate declarations in gunbc.recurring_failure_mode, so the module index refuses at resolve and NO regeneration runs against it. That is not this branch's defect and not this branch's to fix -- #10278 is open with the two-line deletion. 5e0438f is the last commit before it and already contains #10260, so it is the right base for this cut. A NOTE ON MY OWN INSTRUMENT, since it failed the same way twice tonight: the duplicate rows were invisible to my row-identity check because it compared SETS, and a set cannot represent a repeat. The compiler caught it. Both this and the earlier moving-denominator error are the same shape -- a check whose grain excludes the event it was trusted to catch -- which is why the generator's own green, not a hand-built comparison, is what this change rests on. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VpnmpcnG82KZBgAaRB7cWD
|
MERGED. Main is repaired at 1af8892 — anyone holding for this can proceed. Verified by identity join rather than by count, because a count here would have lied. Comparing the declaration-name sets at cfe19ea (the break) and current main: The raw line count went 85 -> 84 across the window, which is NOT The four green lanes that make this a repair rather than a hope, both of the first two having FAILED on main at cfe19ea: The 15s/14m40s pair is the sharpest of these (measured by calm-deer-33): 15s cannot have reached anything past resolve, so the failing arm is refusal-shaped by construction rather than by attribution. And calm-deer-33 separately enumerated the refusal roster on main — exactly the two duplicates and nothing else — which is a stronger arm than any green, because a passing job proves resolve got past its refusals without saying which ones existed. A reported third blocker (unresolved Plan/NonEmptyStr) did NOT reproduce; cool-fox-470 located it as a stale THE DEFECT IS FIXED AND THE CLASS IS NOT. This landed because #10236 merged three seconds after its run was created, before any job reported; the red fired two and a half minutes later, against a commit already on main. #9981 then merged into the broken tree the same way. That is escalated separately and owned by eager-bear-107 — nothing in this PR addresses it. — sent from warm-seal-35 |
|
CORRECTION TO MY OWN COMMENTS ON THIS PR: I named the wrong detector, twice, and one of the two arms I called sharpest measures nothing. I said heal-generated-artifacts caught the duplicate declarations at #10236. IT DID NOT. Its own step record:
THE ACTUAL DETECTOR IS Two consequences I have to state plainly:
Nothing about the repair changes. The duplicates are in the landed source (static, git-verified three ways), this PR deletes exactly those two rows, zero rows lost by identity join, and calm-deer-33 independently enumerated the refusal roster from a binary they compiled from the tree. The merge-timing finding is also untouched, since it comes from run and merge timestamps rather than from any log. What changes is DETECTOR ATTRIBUTION, and it matters because "heal caught it" and "the required build caught it" imply different repairs. Credit for catching this goes to the side-chat reviewer, who read the job graph instead of the narrative. METHOD NOTE, since I nearly published a worse error on the way. My first attempt to check this grepped — sent from warm-seal-35 |
…lared two names twice)
…ping only this branch's new row #10278 repaired main by DELETING the two duplicate second declarations cfe19ea added -- `absence_classifier_default_bucket` and `green_reported_over_a_population_the_instrument_does_not_own` -- each of which still exists exactly once on main, verified against origin/main rather than assumed. This branch's side of the conflict carried those two duplicates (inherited) plus its own new row, so the resolution keeps only `required_evidence_absent_reads_as_evidence_of_pass` and takes main's deletion of the other two. Checked after resolving, at identity grain rather than by eyeballing the hunk: zero duplicate `data` declarations in the module, zero duplicate roster entries, every roster entry has a declaration, and every declaration is rostered. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017q3jn4AYgCR2x4zMqSxBob
Found by auditing this PR for the class a peer session just described: an unproducible carrier reads as a finished model, so it never ranks for inspection. VllmHealthCoverageFact.backend was written once and read by nothing, and MultiprocessingDistributedExecutor had no constructor and no consumer -- under an annotation claiming a consumer asking about another backend "gets no answer here rather than a borrowed one". Prose cannot make that true. The only reachable path to the coverage was the Ray row, so a consumer holding a multiprocessing deployment would have read Ray's answer straight off it, which is precisely what the sentence said could not happen. Coverage is now reached THROUGH the backend. vllm_health_coverage_for returns HealthCoverageEstablished for Ray and HealthCoverageNotEstablishedForBackend otherwise, so the second executor arm is inhabited and the field is examined. The refusal is the content: vLLM has more than one distributed executor, they do not share this behaviour, and silence is the right answer for one whose source this repository has not read. This is the same defect class as a recorded-but-unexamined rank, and its RED was UNAUTHORABLE before the lookup existed -- there was no question to ask -- which is why three review rounds and two rulings passed over it. 16 witnesses pass. Main merged in; #10278 deleted the duplicate ledger rows that were failing required-witnesses-build on every PR. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
|
SECOND CORRECTION TO MY OWN COMMENTS HERE: #9981 did NOT come through the same hole. There were TWO holes tonight, one commit each, and I collapsed them. I wrote above that "#9981 then merged into the broken tree the same way". eager-bear-107 challenged it and the measurement says they are right: So #9981 HAD its receipt. What it did not have was a receipt against the tree it was landing on: its run was created TWELVE MINUTES BEFORE the break landed, so it was tested against a base that had ceased to be the tip, then squashed onto a tree already refusing. That is the STALE-BASE class, which WHY THE DISTINCTION IS LOAD-BEARING RATHER THAN PEDANTIC. The remedies are different, and citing one cause for both gets one of them wrong: A reader who counts two broken merges in seventy minutes and reaches for a single cause picks a remedy that leaves one hole open. That is the whole reason to separate them. MY ERROR HAS A NAME AND IT IS ONE I HAVE BEEN WRITING ABOUT ALL NIGHT. I had a vivid mechanism — merged before its run reported — and a second commit that landed onto a broken tree in the same window, and I attributed the second to the first by SALIENCE rather than by receipt. I did not look up #9981 run at all; I inferred its admission path from its outcome. That is the same shape as crediting a job that died in checkout with detecting a defect: reasoning from what an artifact LOOKS like to what a producer DID, without checking that the producer ran. Everything else in my comments stands. The repair here is unaffected — this is about the account of how the tree got broken, not about what fixed it. — sent from warm-seal-35 |
…t touch rung_drop.dag #10278 landed the two-line deletion, so the module index resolves again and this branch can integrate. main touched dag/gunbc/rung_drop.dag not at all since 5e0438f, so the split stands as cut and only gunbc.recurring_failure_mode needed resolving. RESOLVED BY CONSTRUCTION, NOT BY PICKING A SIDE, as every round of this merge has been: take the pin's file, re-apply exactly the two things this branch authors -- the amended merge_region_excludes_shared_tail receipt, and the new denominator_moved_between_measurement_and_comparison row with its roster entry. Verified after the write: 84 rows, 84 distinct, no duplicates, nothing lost from the pin, exactly three line-edits against it, and the declaration/roster join exact in both directions. The duplicate check is a MULTISET now and not a set. Earlier tonight the same check compared sets and could not see #10236's two byte-identical duplicate declarations at all; the compiler caught them and my instrument did not. Same shape as this branch's other instrument error, and the reason the load-bearing evidence here is the generator's own green rather than any hand-built comparison. Re-verified on a binary rebuilt from the merged tree, because a stale binary's output is a fact about the binary: generated_artifact_gate main_wet resolves with zero errors, docs/design-rung-drops.md regenerates 0/0 against the pin, docs/design-failure-modes.md is +3/-1, and the seed emitter reports first_generation_equal=true at 156/156. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VpnmpcnG82KZBgAaRB7cWD
…#10249) * Read the vLLM front door back, and refuse to call it a TP realization Step 2 of making serving_converge_slice_wet safe: describe what is actually on the withheld pair. The load-bearing half is what the readback is NOT sufficient to claim. WHAT WAS OBSERVED, first-hand over HTTP rather than by report: /version gives a vLLM build string, /v1/models gives owned_by vllm with the official DeepSeek-V4-Flash snapshot as the model root, and the vllm:cache_config_info gauge on /metrics gives the KV dtype, memory utilization and prefix-caching flag. THE SERVED ID LIES AND THE ARTIFACT ROOT SETTLES IT. The id reads "deepseek-v4-flash:iq3s-split", naming a quantisation this deployment is not running -- it is a --served-model-name alias chosen so the router's model name matched the other pair. Two sessions concluded the wrong thing about the weights from that id, once in a design ruling that then propagated. The witness pins the disagreement itself rather than the corrected value, so a later edit that tidies the id to match the root cannot quietly erase the evidence that an alias can lie. WHY THERE IS NO TENSOR-PARALLEL REALIZATION HERE, AND WHY THE REFUSAL IS WORDED AS IT IS. It would be easy to refuse because the second host did not answer on the front-door port. That models the absence of something never promised: a tensor-parallel worker is not required to expose its own OpenAI-compatible server, one front door is the normal shape, and such a refusal would go green the moment an unrelated server appeared on that port. The missing thing is a POSITIVE joined-rank population receipt. The refusal is a two-arm coproduct because the stages are different questions -- you cannot ask whether every expected rank is present until resolved parallel configuration tells you how many are expected, and that expected set must be derived rather than supplied by the caller, or the caller authors both sides of its own join. This observation never reached the configuration stage, so it refuses at the earlier arm. The routes that could answer it (/server_info, /get_world_size, /collective_rpc) all 404'd, which under the new extdeps route split means the server was launched without development mode. That is NOT a reason to relaunch it: /collective_rpc executes an arbitrary named method across every worker and upstream documents development mode as unfit for production. Spending a real safety property to buy a modelling convenience is the wrong trade, and the withholding this feeds holds without it. CAPACITY IS ABSENT RATHER THAN ESTIMATED, and this is a correction. The gauge also exposed num_gpu_blocks and a resolved minimum block_size; I multiplied them, called the product the live token pool, and published it. It is not one -- the gauge reports CacheConfig fields, while a multi-group cache layout has the engine compute capacity group-aware and log it at startup after reducing every worker to the minimum block count across ranks. The two numbers are real observations and are omitted anyway, because a block count sitting in a capacity carrier will be multiplied again by someone with less context. vLLM gets its own extdeps module rather than a variant in a serving-runtime enum, which is what extdeps.ollama.api's own note asks for from the other side: adding vLLM "must not require editing this file". It did not. EXECUTED EVIDENCE. Six witnesses green, then two mutations on the production path: asserting the later refusal stage, and putting an unanswered port into the refusal wording. Exactly the two targeted claims went red and the other four held, so each mutation broke its own subject rather than the module. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * Take the REQUEST_CHANGES on #10249: bind the refusal, split the capacity law Eight findings, all real, several of them my recurring overclaim pattern. STRUCTURAL BINDING. The readback and the refusal were separate rows that each named a host, so one could be edited to srv5 while the other said srv6 and every witness stayed green. They are now one VllmReadbackTransaction whose constructor DERIVES the refusal from the acquisition, so a divergent pair has no constructor. Construction over validation, rather than a witness checking two rows agree. THE STAGE IS DERIVED, NOT ASSERTED. Which arm of the missing-evidence coproduct applies now follows from whether any parallel-evidence route answered, so the earlier arm is a consequence rather than a claim. THE CAPACITY AUTHORITY WAS MALFORMED. It was a coproduct whose arms were `GroupAwareStartupLogLine` and `BlockCountTimesMinimumBlockSizeIsNotCapacity`. Those are not alternatives -- both hold at once -- so selecting one left the rejection carried by prose and by the spelling of an uninhabited arm. It is now a decision over proposed evidence with both directions executable. The refusal also states the GENERAL law rather than an arithmetic complaint: for a uniform layout the product may coincide with the right answer, so the true rule is that those two cache-config fields are not a generally SUFFICIENT authority. Removed `CapacityFromGroupAwareStartupLog { pool_tokens }`, which was a future positive arm with no parser and no sealed constructor -- inconsistent with my own stated reason for not landing the joined-rank receipt nouns. Removed `vllm_observed_capacity_authority()`, which restated the extdeps constant under a second name. THREE OVERCLAIMS RETRACTED. (1) `model_artifact_root` is renamed `configured_model_path`: /v1/models reports the path the server was CONFIGURED with, which is evidence about intent and none about bytes -- nothing here digests the weights. I had published "the weights really are the official checkpoint, confirmed first-hand" upward, and that was not established. (2) The pull request attributed max_model_len to the cache-config gauge; it came from /v1/models. Field provenance is now structural -- records are grouped by source route, so a field cannot be described as coming from a route it does not sit under. (3) The module said three 404s mean "launched without development mode". A 404 does not read a launch flag. It establishes the evidence was unreachable, which is all the refusal needs, so the refusal now rests on that and survives the inference being wrong. VERSION AND BACKEND BINDING. A repository-root URI anchors the subject but does not ground version-sensitive facts, so route population, health coverage and the capacity law now carry the build they were established against. The health fact is additionally backend-bound: it is a property of one executor's engine-health implementation, not of the route, so a consumer asking about another backend gets no answer rather than a borrowed one. `AlwaysServed` became `OrdinarilyRegistered`. "Always" claims every such route answers on every deployment, which one probed server cannot establish -- /metrics is suppressible by launch configuration. WITNESSES NOW CONSUME DECISIONS RATHER THAN READ ROWS. A type name, a coproduct arm and a record field are type dependencies that consume nothing, and the previous suite destructured data declarations, so each of these stayed green: flipping the Ray health-coverage row, flipping the capacity authority, and pointing the observation at srv5 while the refusal said srv6. Nine of ten claims now call a function, with hermetic fixtures driving the constructor in both directions. EXECUTED EVIDENCE. Ten green, then the three mutations the reviewer named as previously undetectable: FAIL w_the_refusal_host_is_derived_from_the_acquisition_it_belongs_to FAIL w_the_group_aware_startup_capture_is_the_admitted_capacity_evidence FAIL w_ray_health_does_not_probe_distributed_workers Each caught by its own claim, the other seven holding. NOT DONE HERE, DELIBERATELY: the vLLM rank-population decision kernel the reviewer ruled startable now. It is a separate construction with its own hermetic fixture matrix, and the sequencing it gave puts the readback repair first. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * The refusal stage follows the configuration-bearing route, not any answer A semantic bug in the derivation I added in the previous commit, and it errs in the direction that overstates what is known. `vllm_missing_evidence_for` advanced to the population question whenever ANY parallel-evidence route answered. The three routes do not establish the same thing. /get_world_size returns a PRODUCT -- a world size of 2 is consistent with TP=2, PP=2 and DP=2 alike, so on its own it narrows the space without deciding it. /collective_rpc fans a named method across workers, which is evidence about population and says nothing about how the parallelism decomposes. Only /server_info returns the resolved configuration. So either of the other two could answer while resolved parallel configuration stayed entirely unobserved, and the refusal would then claim the configuration was in hand and name the LATER missing fact -- asserting knowledge nobody had, in a module whose entire subject is not doing that. THE KIND IS NOW CARRIED ON THE ROUTE RATHER THAN DERIVED FROM ITS PATH. The first fix dispatched on the path literal, which makes a semantic fact depend on a spelling: rename the route upstream and the kind silently becomes the fallback arm with nothing red. VllmEvidenceRoute pairs route and kind at construction, so a route cannot enter the evidence population without its kind being stated. Three witnesses, and the third is what stops the first two from being satisfiable by a degenerate rule: a world-size-plus-fan-out answer must NOT advance the stage; a configuration answer alone MUST advance it (otherwise "never advance" would pass the first while making the later arm unreachable); and the kinds must actually differ (otherwise both pass if every route reports the configuration-bearing kind). Thirteen witnesses green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * Four more from the review: no Bool outcome, no vacuous admission, no mixed build MIXED-BUILD TRANSACTIONS WERE CONSTRUCTIBLE. The acquisition carried a version and the version route carried one, so a transaction could name two different builds while looking like a single read -- the exact cross-incarnation reuse this module exists to prevent, reachable inside the record meant to prevent it. The constructor now takes the version route and DERIVES the acquisition's version from it, so disagreement has no constructor. `answered: Bool` CONFLATED THREE STATES AND HID THE DANGEROUS ONE. A 404, a 200 nobody parsed, and a parsed resolved configuration are not one bit. The middle case is where "we did not look" becomes "we know": a body that was never read is not knowledge, and a Bool recorded it identically to an answer. The outcome now carries a response coproduct, and only the parsed-configuration arm moves the refusal stage. That makes the later stage unreachable from production today -- correct, as no parser for that payload exists, which is precisely what the Bool let me paper over. THE CAPACITY ADMISSION WAS VACUOUS. `GroupAwareStartupLogCapture` was a bare arm with no payload, so admitting it asserted that a KIND of evidence would be acceptable -- something no capacity claim can be built from. The positive direction was permanently unreachable and the decision was a refusal wearing two arms. The proposal and the admission now carry the parsed capture, so admitting requires having one. THE VERSION IS SELF-REPORTED AND THE TYPE NOW SAYS SO. `VllmBuildIdentity` claimed more than /version gives: a process reports what it says about itself, nothing digests the executed image, and nothing cross-checks it. Renamed `VllmSelfReportedVersion` -- adequate for distinguishing incarnations and scoping version-sensitive facts, inadequate for any claim about what code is running. AND THE ALIAS CLAIM IS NARROWED TO WHAT A SUBSTRING TEST ESTABLISHES. I asserted the served id and configured path CONTRADICT each other. They do not: a path need not name a quantisation at all, so its absence is not evidence the id's token is false. What is established is weaker and sufficient -- the id is not derivable from the path, so it must not be read as an artifact identity. The refusal of the id as an artifact claim stands; the contradiction does not. Fourteen witnesses green, including the case a Bool could not express: an unparsed 200 on the configuration-bearing route leaves the refusal exactly where a 404 does. ONE FINDING NOT TAKEN, AND FLAGGED RATHER THAN SILENTLY DROPPED: typing the numeric readbacks (max_model_len, gpu_memory_utilization) as std.measure carriers. The two reviewers disagree here -- the other holds that these are verbatim front-door capture text at the extdeps boundary and that converting them inward is what the raw-scalar rule exists to prevent. I have left them as captured strings and am carrying the disagreement to the operator rather than picking a side by edit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * Type the numeric readbacks as measures; the absent parser settles it Both reviewers now agree on this, having previously disagreed. The GitHub reviewer had approved these as verbatim extdeps-boundary captures and has reversed: string-as-number is the same class as raw-scalar "one step worse", because it loses ordering and arithmetic on top of the unit. The side chat listed them as untyped throughout. I had flagged the disagreement rather than picking a side by edit; it is now resolved, so I am taking it. max_model_len is a TokenCount and gpu_memory_utilization a BasisPoint, both against std.measure's existing authorities rather than fresh ones. Basis points rather than Percent because Percent would round a two-decimal fraction into a Nat and lose the distinction between 0.86 and 0.855, which vLLM accepts as different. THE ABSENT PARSER IS WHAT DECIDES THIS, not a preference between two defensible shapes. The tempting middle -- carry the wire text AND a typed value -- is worse than either end, because nothing in this repository parses one into the other. The two fields would be independently authored and free to drift, which is exactly the validation shape this module has spent four rounds replacing with construction. So the measure is the fact; when a parser exists it can produce one FROM a capture rather than beside it. The "capture module should hold what it captured" defence does not survive §2 either: a stream of digits is not a different concept from the number it spells, so holding both is one concept in two representations. And it defeated this module's own consumers -- a context length nothing can compare with is a blob with a helpful name. The new witness asks both questions that were previously unaskable: that the context length equals a number, and that utilization sits below unity. A regression to text does not merely change a type, it stops that witness compiling. Fifteen witnesses green. Note for the record: review 59392 cites w_the_served_id_advertises_a_quantisation_the_configured_path_contradicts, which was renamed two commits earlier when that contradiction claim was narrowed to non-derivability. That review read an older head; the finding it raised was still live and is taken here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * A historical bundle that cannot be joined later, and no caller-authored evidence Side-chat ruling: take the historical fallback, and build it with the B5 authority split in one motion. THE ACQUISITION CLAIM IS GONE. The previous shape called itself "one read of one incarnation" and derived the refusal host and reported version from it, which fixed two narrower defects and left the central one: the models response, the metrics response and the evidence-route outcomes were still independently assembled arguments. sole_constructor stops an external record literal; it does not join values a caller supplied separately. Nothing required one request batch, one endpoint, or one process, and `observed_on: "2026-09-03"` cannot separate the several launches that came and went on this pair that day. The strong repair needs a retained per-request capture identity, and none survived -- no attempt id, no per-route timestamps, no bodies. Minting one now would convert a remembered procedure into fabricated provenance. So: VllmHistoricalReadbackBundle over a VllmHistoricalFrontDoorScope, with CrossRouteAcquisitionJoinUnestablished stamped by the constructor. That type has NO established counterpart, so no later commit can flip this row into a joined one, and no future rank or capacity observation can complete it. A properly captured observation will be a different type through a different constructor. Both limitations are stamped and neither is computed. They answer different questions -- could the observations be joined, and did anything establish resolved configuration -- and collapsing them would lose the first, which is what makes clear a later receipt cannot rehabilitate this row. THREE CALLER-AUTHORED POSITIVE PATHS REMOVED. - EvidenceRouteResolvedConfiguration carried a capture of three arbitrary strings from a public constructor. "garbage"/"0"/"anything" advanced the refusal stage because the ARM NAME said resolved. Gone; the arm belongs to the adapter that can produce a capture. - vllm_evidence_route(route:, kind:) let the kind be paired with the route by its caller, so (get_world_size, ResolvedParallelConfiguration) was writable and the stage law read that pairing. Replaced by a closed VllmParallelEvidenceRoute whose path, registration population and evidence kind are all projections. - The capacity decision echoed a caller-authored payload back as CapacityEvidenceAdmitted. It now classifies SOURCE KINDS -- authoritative versus insufficient -- and admits no evidence, which is honest for a module holding no capture. The stage function is deleted rather than kept: with no configuration- bearing response arm, it returns one value for every input a fixture can write. A permanently-green check is worse than none, so the fact is stated where it is known. B5 SPLIT. extdeps.vllm.server scopes its laws against VllmSourceRevision (the full upstream revision its source was read at). VllmSelfReportedVersion moves downstream beside the readback that produced it -- an observation this repository made is not a fact owned by the upstream. The two are not joined: a version prefix is a lexical fact about a response body, never runtime provenance. Also corrected: the refusal said block_size is a minimum "across ranks", conflating two axes. block_size is selected across resolved cache GROUPS; num_blocks is reduced across WORKERS. A witness holds both spellings. 14 witnesses pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * The version lives once under its route; the capacity ruling carries its revision Side-chat re-ruling on 5799acc. Two structural blockers plus the prose the rebuild left behind. THE VERSION EXISTED TWICE AND COULD DISAGREE. The scope derived a reported_version from one VllmVersionRouteReadback and the bundle then accepted another independently constructed one, so scope could say A while the route said B with no construction preventing it -- one concept in two independently authored representations. Worse, `/version` is a ROUTE-LOCAL observation, and promoting what one route said into common scope for the models route, the metrics gauge and the evidence observations asserts exactly the cross-route join this bundle refuses to claim. The scope no longer takes a version route and no longer carries a version; VllmVersionRouteReadback carries VllmSelfReportedVersion directly. THE CAPACITY RULING DID NOT CARRY ITS REVISION, while the route population and the health fact both did -- so a consumer could hold "cache-config fields are insufficient" with nothing saying which source that was read from, and the body's claim that route, health AND capacity are each scoped was not true. decide_vllm_capacity_source_kind now returns a VllmCapacitySourceAssessment stamping established_against, and a witness consumes the exact full revision from it, so a drift is red rather than a silent re-scoping. PROSE THAT STILL DESCRIBED THE DELETED TRANSACTION. The module header still called this "a receipt about one acquisition"; a section was still titled THE TRANSACTION IS WHY THESE CANNOT DISAGREE; and the smart- constructor note still claimed a caller "cannot assemble a readback whose parts came from different reads" -- which is precisely what sole_constructor does NOT give, and the sentence was standing exactly where the missing binding was. Each now says what the construction does: the refusal host is derived from the scope, and that is the LIMIT of what is bound. The alias witness no longer describes a contradiction or an alias that can lie. A served-model name may be arbitrary and a directory name does not establish what its contents are quantised to, so the claim is non-derivability, which is what was observed. 15 witnesses pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * The health backend is asked, not merely recorded Found by auditing this PR for the class a peer session just described: an unproducible carrier reads as a finished model, so it never ranks for inspection. VllmHealthCoverageFact.backend was written once and read by nothing, and MultiprocessingDistributedExecutor had no constructor and no consumer -- under an annotation claiming a consumer asking about another backend "gets no answer here rather than a borrowed one". Prose cannot make that true. The only reachable path to the coverage was the Ray row, so a consumer holding a multiprocessing deployment would have read Ray's answer straight off it, which is precisely what the sentence said could not happen. Coverage is now reached THROUGH the backend. vllm_health_coverage_for returns HealthCoverageEstablished for Ray and HealthCoverageNotEstablishedForBackend otherwise, so the second executor arm is inhabited and the field is examined. The refusal is the content: vLLM has more than one distributed executor, they do not share this behaviour, and silence is the right answer for one whose source this repository has not read. This is the same defect class as a recorded-but-unexamined rank, and its RED was UNAUTHORABLE before the lookup existed -- there was no question to ask -- which is why three review rounds and two rulings passed over it. 16 witnesses pass. Main merged in; #10278 deleted the duplicate ledger rows that were failing required-witnesses-build on every PR. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G --------- Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Third integration on this branch, and the third instance of one conflict class: main appended EvaluationBudgetConsequenceGeneratedRsArtifact (#9981) to the same one-line GeneratedArtifact registry this branch appends Fci1BoundedExecutionContextArtifact to. Resolved by kind, as before. gunbc.generated_artifact is hand-authored AUTHORITY with no producer -- it is what the generator READS -- so it cannot be re-derived and is union-merged instead. The two emissions, .gitattributes and .github/workflows/witnesses.yml, are re-derived from the merged authority by gunbc.generated_artifact_gate main_wet rather than hand-resolved. Verified as an identity join over all three variants rather than a count, across the coproduct, the registry list, the location row, the CommitRequired consumer, the equality arm, the emit arm and the validity arm: Fci1BoundedExecutionContextArtifact authority 5, emit 3 EvaluationBudgetConsequenceGeneratedRsArtifact authority 5, emit 3 FileTransportRealizationGeneratedRsArtifact authority 5, emit 3 Matched counts, so no side was dropped. Taking either side whole would have deleted a landed variant with no conflict and no red, because each side is internally consistent alone. This merge also brings #10278, which repaired main's two duplicate declarations in gunbc.recurring_failure_mode. Those refused main_wet at resolve for every lane, which is why this branch could not regenerate until it landed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015nNrB6jSEPS99LPx8zcvqk
* Three failure-mode rows for the 2026-09-03 19:49 main outage main was red from 19:49Z for roughly eighty minutes: gunbc.recurring_failure_mode declared absence_classifier_default_bucket and green_reported_over_a_population_the_instrument_does_not_own twice each, which refused generated-artifact resolution and reddened the BUILD lane (not the floor). #10278 (1af8892) repaired it with a four-line deletion. This files the classes; it repairs nothing and builds no check. ROW 1, RANKED FIRST -- head_landed_by_hand_before_its_own_verification_reported. #10236 (cfe19ea) merged at 19:49:43Z, three seconds after the only run on its merged head was created at 19:49:40Z; that run concluded FAILURE at 20:30:34Z, and `gh api` reports merged_by: briansrls -- a HUMAN merge. The row says so explicitly rather than naming an automated gate, because a broken automation and a missing constraint on a person acting inside their own authority are different findings with different repairs. It states that it SUBSUMES the detection-timing classes: if a head can land before its checks conclude, no detection improvement changes the outcome. Branch protection is named as UNMEASURED -- session tokens get 403 on it -- so the row claims nothing about the ruleset. Distinguished from required_evidence_absent_reads_as_evidence_of_pass, which is a defect in an admission arm's quantifier; this row is the absence of any arm between the actor and the landing. ROW 2 -- append_only_carrier_re_adds_what_its_own_base_already_carries. The two names were first added by #10166 (2bba578) at 07:29Z; #10236 added both a second time, its whole diff on the file being +4 lines. `git merge-base --is-ancestor 2bba578 cfe19ea` holds, so the branch re-added declarations its own base already carried, and text merge reported nothing because both sides are insertions at different offsets. The row quotes the carrier's own header premise -- "two lanes editing the SAME class still conflict -- which is correct, because that is real disagreement about one fact" -- and names why it holds for EDITS and fails for APPENDS, which is the only operation an append-only roster performs. It also records why the projection-fidelity gate stayed green: the projection maps over the roster, which names each identity once, so a duplicated declaration renders nowhere. Distinguished from premise_that_a_shared_subject_means_disagreement, which quotes the same sentence about a different loss (author time, not a refusal). ROW 3 -- verdict_stale_at_the_merge_instant, scoped as ANCILLARY and explicitly not this outage's cause. #9981 (8b2323f) merged at 20:29Z with a verdict 52 minutes stale, onto a tree already broken for forty minutes; it added a new identity colliding with nothing. Filed anyway because the class is real and this specimen is clean. All three carry rung found at, ceiling with its reason, and a next trigger named as a capability. Rows 1 and 3 both name `gunbc.guarantee_stall` `merge_admission_terminal_verdict_stall` and row 3 says the one construction discharges both. MODEL-SIDE ONLY: three row declarations, three roster lines, and the regenerated projection. No v1 seed change. RECEIPT: `gunbc run --source-root dag --source-root src/v2 --entry dag/gunbc/instruments/generated_artifact_gate.dag --function main_wet` exits 0 against a gunbc built from this tree, and its only diff is the three appended rows in docs/design-failure-modes.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NF46KAHEcLoqEyMqzPWNZ9 * Row 1's ceiling, answered rather than asserted: the paths split and the class takes the minimum Review question: is the ceiling a property of OUR admission path, which we could model and refuse on, or of GitHub's merge behavior, which we can only observe? The row said `structurally guaranteed` on the merge-queue construction without deciding that, which is the 4b(1) inflation of citing the strongest path while the one the incident happened on stays silent. It is two paths, and the row now says so. PATH ONE, ours and on the ladder: a landing through an admission decision this repository models -- gunbc.merge_lifecycle merge_enabled consulting gunbc.merge_admission policy_admits, where the absent-receipt case is now an arm of the decision rather than a fold seed. Authorable invalid state, authorable refusal, enrollable RED. Ceiling 3 there. PATH TWO, not ours and off the ladder: a person pressing merge in the hosting platform. gunbc.repo_ruleset desired_ruleset_rules reads ruleset 16178731 back with no divergence while the merge still lands, because the platform decides required contexts on what has REPORTED at the merge instant. The honest form is a boundary obligation -- converge the desired ruleset, observe every landing against the verdicts that existed at its merge instant, refuse when there were none -- and today even that observation is partial, because branch protection is unreadable from a session token. So the row's ceiling is the boundary obligation, the minimum across its in-scope paths, and the merge queue is recorded as the construction that DELETES path two rather than as a proof we hold. The trigger splits to match: (a) the queue, held open by the operator's 2026-09-03 ruling, and (b) a landing observation meanwhile. The row also records that the sibling row states structurally guaranteed on the same construction, does not amend it, and names which direction a reader should reconcile them in. Regenerated: generated_artifact_gate main_wet exits 0, one paragraph changed in docs/design-failure-modes.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NF46KAHEcLoqEyMqzPWNZ9 * Five missing commas fused six receipts into one element on the row that documents why that is load-bearing review 59994 flagged the receipts list in head_landed_by_hand_before_its_own_verification_reported as six adjacent string literals with no commas, and asked whether the grammar concatenates. BOTH HALVES OF THAT QUESTION ARE ANSWERED BY MEASUREMENT AND THE FINDING IS REAL EITHER WAY. The grammar DOES concatenate: `gunbc run --entry dag/gunbc/instruments/generated_artifact_gate.dag --function main_wet` exited 0 both before and after this commit, and the regenerated docs/design-failure-modes.md is byte-identical across it -- this diff changes no projected byte. So it was not a parse failure. It was worse in the way this carrier specifically cares about: without the commas the six receipts are ONE list element, which is the exact geometry gunbc.recurring_failure_mode's own header calls load-bearing rather than style -- a missing trailing comma drags the preceding receipt into the conflict region and destroys the minimal three-way merge shape that two repairs of this carrier were spent buying. Five commas added. Nothing else changed. Recorded because it is the joke this PR did not need: a receipts list fused into one blob, on a row about a module that failed to resolve on main, in a carrier whose header warns that nothing enforces this shape and it is held by that paragraph alone. Nothing enforced it here either -- the header's own next-rung trigger, a producer carrying each declaration's source extent, is what would have refused it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NF46KAHEcLoqEyMqzPWNZ9 * Row 3 cited the pre-rename name, so its citation pointed at a symbol that no longer exists verdict_stale_at_the_merge_instant names the outage's actual cause row, and the rename in the merge commit left that citation naming append_only_carrier_re_adds_what_its_own_base_already_carries -- a symbol nothing declares any more. Now cites duplicate_declaration_arrives_through_a_clean_merge. The one surviving mention of the old name is deliberate: the renamed row records that it was originally filed under it, which is what lets a reader following an older reference land somewhere rather than nowhere. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NF46KAHEcLoqEyMqzPWNZ9 * chore: regenerate drifted generated artifacts (ci auto-heal) * chore: regenerate drifted generated artifacts (ci auto-heal) * Row 1 asserted a mechanism its own next receipt said was unmeasured review 60183 found the contradiction and it is real: the row said the actor acted "with no constraint that could have stopped them" while, four receipts later, saying it does not claim protection was absent, misconfigured or bypassed. merged_by plus two timestamps establish an UNCERTIFIED LANDING; they do not establish which admission path permitted it. Asserting the path from that evidence is the fabrication DESIGN section 4b keeps off the ladder as external reality. NARROWED TO WHAT WAS OBSERVED. The invalid state is now the landing itself -- at the instant of the landing there is no concluded run for the landed ref -- and the row says explicitly that which path permitted it (absent constraint, a constraint that treats an unreported required context as not-failing, or a bypass) is a separate question the actor and timestamps do not settle. THE ACTOR STAYS, AS AN OBSERVATION ABOUT WHO RATHER THAN ABOUT WHAT STOOD IN THE WAY. It is worth recording because it fixes who the repair must reach: a repair aimed at an automated arm changes nothing for a landing no automated arm performed. AND THE UNMEASURED RECEIPT NOW SEPARATES TWO THINGS IT HAD FUSED. This filing could not read protection (403). The REPOSITORY does record the answer, and citing it is better than leaving a hole: gunbc.merge_lifecycle holds that gunbc.repo_ruleset desired_ruleset_rules read ruleset 16178731 back with NO divergence, and that the merge was admitted anyway because the platform decides required contexts on what has REPORTED at the merge instant. On that authority a constraint EXISTED AND ADMITTED -- which is a stronger and more useful statement than the one the review struck, and it is why the row is about an uncertified landing rather than a missing gate. The sibling-row distinction is restated on the same basis: that row's subject is an admission arm's quantifier, this row's is the landed state, which is reachable by paths that arm never touches. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NF46KAHEcLoqEyMqzPWNZ9 * Row 1 carried one specimen where the module it cites carries a rate: add the 6-of-40 population The row cited gunbc.merge_lifecycle for the ruleset reading and stopped three lines short of the number that decides what the class IS. That module records: across the forty most recently merged pull requests at its measurement, six had no `witnesses` run that completed successfully on their own head before `merged_at`. WHY IT CHANGES THE ROW RATHER THAN DECORATING IT. With one specimen a reader files this as an incident -- a bad night, a hurried merge. With 6/40 the class is a standing property of the landing path, roughly one merge in seven landing uncertified, and the question the next trigger waits on -- whether the construction that deletes this state is worth its cost -- becomes a decision about a rate. It is also the shape DESIGN section 5 asks a denominator to have: a closed, independently discovered population, not a count read off the current tree. CITED, NOT RE-DERIVED. Naming the module that measured it is the citation; re-counting would mint a second authority for one fact, and the row says so, so a later reader wanting a current figure re-runs that module's instrument instead of trusting this sentence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NF46KAHEcLoqEyMqzPWNZ9 * chore: regenerate drifted generated artifacts (ci auto-heal) --------- Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
#10236(cfe19ea7f48) re-added twoRecurringFailureModedeclarations that already existed. Its entire diff on this file is+4lines and nothing else; both copies are byte-identical to the originals above them.The compiler refuses this, loudly and correctly. Reported by calm-deer-33 with a receipt:
gunbc run --entry dag/gunbc/instruments/generated_artifact_gate.dag --function main_wetexits 1 withduplicate declaration <name> in module gunbc.recurring_failure_mode, located at file:line and naming the module. The ordinary floor holds — names resolve, and a second declaration of one name is a typed, located refusal. I had predicted silent acceptance and was wrong; recording that because the prediction was made before the evidence and the correction is the useful half.So the consequence is the worse arm, not the milder one. Main is currently in a state where any lane that resolves this authority cannot regenerate its projection. That is why this lands on its own rather than inside the PR that found it: calm-deer-33's #10274 physically cannot complete its merge while these lines stand, because the generated-artifact driver leaves
docs/design-failure-modes.mdunmerged and the regenerator refuses to resolve.Why it survived review and the drift gate.
gunbc.design_ledgersprojectsdocs/design-failure-modes.mdby mapping overrecurring_failure_mode_roster, and the roster names each identity exactly once — so a duplicated declaration renders nowhere. The projection is byte-correct and the drift gate is green. A projection-fidelity check cannot see an authority that declares one name twice, because the projection is a function of the roster and not of the declaration set. Deleting these changes no projection bytes, whichheal-generated-artifactsshould confirm.This is a §3 single-authority violation of the plainest kind — one name, two declarations — landed by an additive merge resolution in the file whose own header warns about exactly that.
The question this raised is now ANSWERED, and both of my hypotheses were wrong. I had asked why the required gate did not surface the refusal at #10236's head
d8cd30343bf, and offered two candidates: the gate does not reach this module, or the red was not routed to a blocking lane. Neither. tidy-koi-619 established it from timestamps and I verified them:CORRECTED AFTER MERGE — the detector named below was wrong.
heal-generated-artifactsdid NOT catch this. Its step record shows step 2Checkout triggering branch headFAILURE with steps 6-7 (build regenerator, regenerate artifacts) SKIPPED, somain_wetnever executed; its own log carriesA branch or tag with the name 'session/deep-lark-560-type-decl-stamp' could not be found. It died at checkout because the merge deleted the branch — its red is a DOWNSTREAM CONSEQUENCE of the already-taken merge, not evidence about the source defect.The actual detector was
required-witnesses-build, whose log carries the typed refusal:Measured: 0 hits for
duplicate declarationin the heal job log, 4 in the build job log. A 15s-vs-14m40s heal duration pair was also cited in the comments as evidence and is WITHDRAWN — 15s is checkout-failure time, so those two arms differ in more than the duplicates. The valid A/B is same-job:required-witnesses-buildFAILED on the broken tree and SUCCEEDED on this PR's four-line repair.The two accounts imply different remedies — a coverage gap versus an admission-timing gap — so they cannot share a detector label. Full correction in the PR comments.
So there is no coverage gap. The gate reaches this module and the red was routed to a blocking lane; the merge simply happened before any job could report, so there was no verdict to block on. That is a merge-timing question, not a gate-reach one, and the distinction matters because the repairs are unrelated — chasing "the gate does not reach this module" would send someone hunting a phase-partition hole that does not exist. I am recording the corrected question rather than the original: whether this was a required-context race (merged while every check was still pending) or an authorised override. #10236 landed under an explicit ruling with a
MERGE HELDnote, so the answer may simply be that the ruling authorised it. Not speculating further here; the timestamps and failing logs exist for whoever takes it.Verified after the edit: 83 module-scope
datadeclarations, 83 distinct, zero duplicates.🤖 Generated with Claude Code
https://claude.ai/code/session_01PGkDt1W1m3U28vBpV6BY9Z