Skip to content

Read the vLLM front door back, and refuse to call it a TP realization - #10249

Merged
briansrls merged 9 commits into
mainfrom
session/eager-pike-541-vllm-observed
Sep 3, 2026
Merged

briansrls merged 9 commits into
mainfrom
session/eager-pike-541-vllm-observed

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Step 2 of making serving_converge_slice_wet safe: describe what is actually on the pair #10225 withholds. The load-bearing half is what the readback is not sufficient to claim.

This body was rewritten. Its first version repeated three claims that were later retracted — that the readback confirmed the official weights, that max_model_len came from the metrics gauge, and that three 404s proved the server was launched without development mode. None of those is true. They are listed under Retractions below rather than silently deleted, because the PR body is what a merger reads and quietly dropping a published claim leaves no trace that it was made.

What was observed

First-hand over HTTP, not by report:

route what it gave
/version a vLLM build string
/v1/models owned_by: vllm, a configured model path, max_model_len 262144
/metrics (vllm:cache_config_info) cache_dtype fp8_ds_mla, gpu_memory_utilization 0.86, prefix caching on

Records are grouped by the route they came from, so provenance is structural: a field cannot be described as coming from a route it does not sit under.

The served alias and the configured path are two separate facts. The id reads deepseek-v4-flash:iq3s-split and carries a quantisation token that is not derivable from the configured path — a --served-model-name alias chosen so the router's model name matched the other pair. That is enough to refuse the id as artifact identity. It is not enough to conclude the alias is false: a served-model name is allowed to be arbitrary, and a directory name does not establish what its contents are quantised to. Two sessions concluded the wrong thing about the weights from that id, once in a design ruling that propagated; the witness pins the non-derivability, which is the part that was observed.

Why there is no TP realization here

Refusing because the second host did not answer on the front-door port would model the absence of something never promised: a tensor-parallel worker is not required to expose its own OpenAI-compatible server, and such a refusal would go green the moment an unrelated server appeared on that port.

What is missing is a positive joined-rank population receipt. The refusal is two-stage because the questions differ: you cannot ask whether every expected rank is present until resolved parallel configuration says how many are expected, and that expected set must be derived, not supplied by the caller.

The route identity decides its kind; neither is supplied by a caller. /get_world_size returns a product consistent with TP=2, PP=2 and DP=2 alike; /collective_rpc reports population without saying how parallelism decomposes. Only /server_info returns resolved configuration. An earlier version paired route and kind at construction, which left the relation caller-authored — (get_world_size, ResolvedParallelConfiguration) was writable, and the stage law read that pairing. It is now a projection of a closed VllmParallelEvidenceRoute, so no combination can be authored.

There is no parsed-configuration arm, and its absence is the point. The previous EvidenceRouteResolvedConfiguration { capture } was built from three arbitrary strings by a public constructor: a caller could write "garbage", "0", "anything", wrap it, and advance the refusal stage because the arm name said the configuration had been resolved. That arm belongs to the adapter that can actually produce a capture.

All three evidence routes 404'd. That is not a reason to relaunch the server. /collective_rpc executes an arbitrary named method across every worker and upstream documents development mode as unfit for production. The withholding this feeds holds without it.

Retractions

  • "The weights really are the official checkpoint, confirmed first-hand." /v1/models reports the path the server was configured with — evidence about intent, none about bytes. Nothing here digests the weights. The field is now configured_model_path.
  • max_model_len attributed to the cache-config gauge. It came from /v1/models.
  • "Three 404s mean the server was launched without development mode." A 404 does not read a launch flag. It establishes the evidence route was unreachable — which is all the refusal needs, so the refusal survives the inference being wrong.
  • A live pool of 371,056 tokens (num_gpu_blocks × block_size). Not this build's capacity law, and the two factors are not even on the same axis: block_size is selected as the minimum across resolved cache groups while num_blocks is separately reduced across workers, so the product joins values from different axes — and a multi-group layout has the engine compute capacity group-aware and log it at startup. The bundle carries no capacity field at all: a block count in a capacity carrier gets multiplied again by the next reader.

Structure

  • This is a historical bundle, not a transaction. The previous shape called itself one read of one incarnation and derived the refusal host from the acquisition — which fixed two narrower defects and left the central one: the models response, the metrics response and the evidence-route outcomes were still independently assembled arguments. sole_constructor stops an external record literal; it does not join values a caller supplied separately, and observed_on: "2026-09-03" cannot separate the several launches that came and went on this pair that day.
  • The cross-route join was never retained, and that is permanent. No attempt id, no per-route timestamps, no captured bodies survived, so minting an acquisition identity now would convert a remembered procedure into fabricated provenance. CrossRouteAcquisitionJoinUnestablished is stamped by the constructor and has no established counterpart — no later commit can flip this row into a joined one, and no future rank or capacity observation can complete it. A properly captured observation will be a different type through a different constructor.
  • Both limitations are stamped, and they answer different questions. Could the observations be joined as one acquisition (no), and did any retained observation establish resolved parallel configuration (no). The second explains the refusal stage; the first is what makes clear a later receipt cannot rehabilitate this bundle.
  • The capacity law classifies source kinds and admits no evidence. It previously echoed a caller-authored two-string payload back as CapacityEvidenceAdmitted. This module holds no capture, so it says which kind of source would be authoritative and stops there.
  • The scoping authority is the upstream source revision. extdeps.vllm.server scopes its route, health and capacity laws against the full revision its source was read at; VllmSelfReportedVersion moved downstream beside the readback that produced it, because an observation this repository made is not a fact owned by the upstream. The two are not joined — a version prefix is a lexical fact about a response body, never runtime provenance.
  • Route population, executor health and the capacity law each carry that revision; health is additionally bound to RayDistributedExecutor, whose engine-health implementation returns without contacting workers.
  • OrdinarilyRegistered, not AlwaysServed — one probed server cannot establish that a route answers on every deployment.

Executed evidence

Fourteen witnesses green, driving constructors and decisions with hermetic fixtures rather than destructuring the pinned row. Both capacity directions execute; the group-versus-rank axes are asserted separately; the stage-derivation function was deleted rather than kept, because with no configuration-bearing response arm it returns one value for every input a fixture can write, and a permanently-green check is worse than none.

Freshness

The front door answered on 2026-09-03 and had stopped answering within the hour; .225/.226 are now silent while production .232/.233 are up. This is a bundle of route-local observations associated with one front-door investigation — not one acquisition and not one identified incarnation, neither of which the retained evidence establishes. It supports what each route reported. It does not support that all values belonged to one engine process, that all requests preceded any restart, or that the self-reported version identifies executed bytes.

🤖 Generated with Claude Code

https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G

gunbc-ci-auto-heal and others added 5 commits September 3, 2026 16:27
Step 2 of making serving_converge_slice_wet safe: describe what is actually on
the withheld pair. The load-bearing half is what the readback is NOT sufficient
to claim.

WHAT WAS OBSERVED, first-hand over HTTP rather than by report: /version gives a
vLLM build string, /v1/models gives owned_by vllm with the official
DeepSeek-V4-Flash snapshot as the model root, and the vllm:cache_config_info
gauge on /metrics gives the KV dtype, memory utilization and prefix-caching flag.

THE SERVED ID LIES AND THE ARTIFACT ROOT SETTLES IT. The id reads
"deepseek-v4-flash:iq3s-split", naming a quantisation this deployment is not
running -- it is a --served-model-name alias chosen so the router's model name
matched the other pair. Two sessions concluded the wrong thing about the weights
from that id, once in a design ruling that then propagated. The witness pins the
disagreement itself rather than the corrected value, so a later edit that tidies
the id to match the root cannot quietly erase the evidence that an alias can lie.

WHY THERE IS NO TENSOR-PARALLEL REALIZATION HERE, AND WHY THE REFUSAL IS WORDED
AS IT IS. It would be easy to refuse because the second host did not answer on
the front-door port. That models the absence of something never promised: a
tensor-parallel worker is not required to expose its own OpenAI-compatible
server, one front door is the normal shape, and such a refusal would go green the
moment an unrelated server appeared on that port. The missing thing is a POSITIVE
joined-rank population receipt. The refusal is a two-arm coproduct because the
stages are different questions -- you cannot ask whether every expected rank is
present until resolved parallel configuration tells you how many are expected,
and that expected set must be derived rather than supplied by the caller, or the
caller authors both sides of its own join. This observation never reached the
configuration stage, so it refuses at the earlier arm.

The routes that could answer it (/server_info, /get_world_size, /collective_rpc)
all 404'd, which under the new extdeps route split means the server was launched
without development mode. That is NOT a reason to relaunch it: /collective_rpc
executes an arbitrary named method across every worker and upstream documents
development mode as unfit for production. Spending a real safety property to buy
a modelling convenience is the wrong trade, and the withholding this feeds holds
without it.

CAPACITY IS ABSENT RATHER THAN ESTIMATED, and this is a correction. The gauge
also exposed num_gpu_blocks and a resolved minimum block_size; I multiplied them,
called the product the live token pool, and published it. It is not one -- the
gauge reports CacheConfig fields, while a multi-group cache layout has the engine
compute capacity group-aware and log it at startup after reducing every worker to
the minimum block count across ranks. The two numbers are real observations and
are omitted anyway, because a block count sitting in a capacity carrier will be
multiplied again by someone with less context.

vLLM gets its own extdeps module rather than a variant in a serving-runtime enum,
which is what extdeps.ollama.api's own note asks for from the other side: adding
vLLM "must not require editing this file". It did not.

EXECUTED EVIDENCE. Six witnesses green, then two mutations on the production
path: asserting the later refusal stage, and putting an unanswered port into the
refusal wording. Exactly the two targeted claims went red and the other four
held, so each mutation broke its own subject rather than the module.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
…ity law

Eight findings, all real, several of them my recurring overclaim pattern.

STRUCTURAL BINDING. The readback and the refusal were separate rows that each
named a host, so one could be edited to srv5 while the other said srv6 and every
witness stayed green. They are now one VllmReadbackTransaction whose constructor
DERIVES the refusal from the acquisition, so a divergent pair has no constructor.
Construction over validation, rather than a witness checking two rows agree.

THE STAGE IS DERIVED, NOT ASSERTED. Which arm of the missing-evidence coproduct
applies now follows from whether any parallel-evidence route answered, so the
earlier arm is a consequence rather than a claim.

THE CAPACITY AUTHORITY WAS MALFORMED. It was a coproduct whose arms were
`GroupAwareStartupLogLine` and `BlockCountTimesMinimumBlockSizeIsNotCapacity`.
Those are not alternatives -- both hold at once -- so selecting one left the
rejection carried by prose and by the spelling of an uninhabited arm. It is now a
decision over proposed evidence with both directions executable. The refusal also
states the GENERAL law rather than an arithmetic complaint: for a uniform layout
the product may coincide with the right answer, so the true rule is that those two
cache-config fields are not a generally SUFFICIENT authority.

Removed `CapacityFromGroupAwareStartupLog { pool_tokens }`, which was a future
positive arm with no parser and no sealed constructor -- inconsistent with my own
stated reason for not landing the joined-rank receipt nouns. Removed
`vllm_observed_capacity_authority()`, which restated the extdeps constant under a
second name.

THREE OVERCLAIMS RETRACTED. (1) `model_artifact_root` is renamed
`configured_model_path`: /v1/models reports the path the server was CONFIGURED
with, which is evidence about intent and none about bytes -- nothing here digests
the weights. I had published "the weights really are the official checkpoint,
confirmed first-hand" upward, and that was not established. (2) The pull request
attributed max_model_len to the cache-config gauge; it came from /v1/models. Field
provenance is now structural -- records are grouped by source route, so a field
cannot be described as coming from a route it does not sit under. (3) The module
said three 404s mean "launched without development mode". A 404 does not read a
launch flag. It establishes the evidence was unreachable, which is all the refusal
needs, so the refusal now rests on that and survives the inference being wrong.

VERSION AND BACKEND BINDING. A repository-root URI anchors the subject but does
not ground version-sensitive facts, so route population, health coverage and the
capacity law now carry the build they were established against. The health fact is
additionally backend-bound: it is a property of one executor's engine-health
implementation, not of the route, so a consumer asking about another backend gets
no answer rather than a borrowed one.

`AlwaysServed` became `OrdinarilyRegistered`. "Always" claims every such route
answers on every deployment, which one probed server cannot establish -- /metrics
is suppressible by launch configuration.

WITNESSES NOW CONSUME DECISIONS RATHER THAN READ ROWS. A type name, a coproduct
arm and a record field are type dependencies that consume nothing, and the previous
suite destructured data declarations, so each of these stayed green: flipping the
Ray health-coverage row, flipping the capacity authority, and pointing the
observation at srv5 while the refusal said srv6. Nine of ten claims now call a
function, with hermetic fixtures driving the constructor in both directions.

EXECUTED EVIDENCE. Ten green, then the three mutations the reviewer named as
previously undetectable:

  FAIL w_the_refusal_host_is_derived_from_the_acquisition_it_belongs_to
  FAIL w_the_group_aware_startup_capture_is_the_admitted_capacity_evidence
  FAIL w_ray_health_does_not_probe_distributed_workers

Each caught by its own claim, the other seven holding.

NOT DONE HERE, DELIBERATELY: the vLLM rank-population decision kernel the reviewer
ruled startable now. It is a separate construction with its own hermetic fixture
matrix, and the sequencing it gave puts the readback repair first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
…swer

A semantic bug in the derivation I added in the previous commit, and it errs in
the direction that overstates what is known.

`vllm_missing_evidence_for` advanced to the population question whenever ANY
parallel-evidence route answered. The three routes do not establish the same
thing. /get_world_size returns a PRODUCT -- a world size of 2 is consistent with
TP=2, PP=2 and DP=2 alike, so on its own it narrows the space without deciding it.
/collective_rpc fans a named method across workers, which is evidence about
population and says nothing about how the parallelism decomposes. Only
/server_info returns the resolved configuration.

So either of the other two could answer while resolved parallel configuration
stayed entirely unobserved, and the refusal would then claim the configuration was
in hand and name the LATER missing fact -- asserting knowledge nobody had, in a
module whose entire subject is not doing that.

THE KIND IS NOW CARRIED ON THE ROUTE RATHER THAN DERIVED FROM ITS PATH. The first
fix dispatched on the path literal, which makes a semantic fact depend on a
spelling: rename the route upstream and the kind silently becomes the fallback arm
with nothing red. VllmEvidenceRoute pairs route and kind at construction, so a
route cannot enter the evidence population without its kind being stated.

Three witnesses, and the third is what stops the first two from being satisfiable
by a degenerate rule: a world-size-plus-fan-out answer must NOT advance the stage;
a configuration answer alone MUST advance it (otherwise "never advance" would pass
the first while making the later arm unreachable); and the kinds must actually
differ (otherwise both pass if every route reports the configuration-bearing kind).

Thirteen witnesses green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
…mixed build

MIXED-BUILD TRANSACTIONS WERE CONSTRUCTIBLE. The acquisition carried a version and
the version route carried one, so a transaction could name two different builds
while looking like a single read -- the exact cross-incarnation reuse this module
exists to prevent, reachable inside the record meant to prevent it. The
constructor now takes the version route and DERIVES the acquisition's version from
it, so disagreement has no constructor.

`answered: Bool` CONFLATED THREE STATES AND HID THE DANGEROUS ONE. A 404, a 200
nobody parsed, and a parsed resolved configuration are not one bit. The middle case
is where "we did not look" becomes "we know": a body that was never read is not
knowledge, and a Bool recorded it identically to an answer. The outcome now carries
a response coproduct, and only the parsed-configuration arm moves the refusal
stage. That makes the later stage unreachable from production today -- correct, as
no parser for that payload exists, which is precisely what the Bool let me paper
over.

THE CAPACITY ADMISSION WAS VACUOUS. `GroupAwareStartupLogCapture` was a bare arm
with no payload, so admitting it asserted that a KIND of evidence would be
acceptable -- something no capacity claim can be built from. The positive direction
was permanently unreachable and the decision was a refusal wearing two arms. The
proposal and the admission now carry the parsed capture, so admitting requires
having one.

THE VERSION IS SELF-REPORTED AND THE TYPE NOW SAYS SO. `VllmBuildIdentity` claimed
more than /version gives: a process reports what it says about itself, nothing
digests the executed image, and nothing cross-checks it. Renamed
`VllmSelfReportedVersion` -- adequate for distinguishing incarnations and scoping
version-sensitive facts, inadequate for any claim about what code is running.

AND THE ALIAS CLAIM IS NARROWED TO WHAT A SUBSTRING TEST ESTABLISHES. I asserted
the served id and configured path CONTRADICT each other. They do not: a path need
not name a quantisation at all, so its absence is not evidence the id's token is
false. What is established is weaker and sufficient -- the id is not derivable from
the path, so it must not be read as an artifact identity. The refusal of the id as
an artifact claim stands; the contradiction does not.

Fourteen witnesses green, including the case a Bool could not express: an unparsed
200 on the configuration-bearing route leaves the refusal exactly where a 404 does.

ONE FINDING NOT TAKEN, AND FLAGGED RATHER THAN SILENTLY DROPPED: typing the numeric
readbacks (max_model_len, gpu_memory_utilization) as std.measure carriers. The two
reviewers disagree here -- the other holds that these are verbatim front-door
capture text at the extdeps boundary and that converting them inward is what the
raw-scalar rule exists to prevent. I have left them as captured strings and am
carrying the disagreement to the operator rather than picking a side by edit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
Both reviewers now agree on this, having previously disagreed. The GitHub reviewer
had approved these as verbatim extdeps-boundary captures and has reversed:
string-as-number is the same class as raw-scalar "one step worse", because it loses
ordering and arithmetic on top of the unit. The side chat listed them as untyped
throughout. I had flagged the disagreement rather than picking a side by edit; it
is now resolved, so I am taking it.

max_model_len is a TokenCount and gpu_memory_utilization a BasisPoint, both against
std.measure's existing authorities rather than fresh ones. Basis points rather than
Percent because Percent would round a two-decimal fraction into a Nat and lose the
distinction between 0.86 and 0.855, which vLLM accepts as different.

THE ABSENT PARSER IS WHAT DECIDES THIS, not a preference between two defensible
shapes. The tempting middle -- carry the wire text AND a typed value -- is worse
than either end, because nothing in this repository parses one into the other. The
two fields would be independently authored and free to drift, which is exactly the
validation shape this module has spent four rounds replacing with construction. So
the measure is the fact; when a parser exists it can produce one FROM a capture
rather than beside it.

The "capture module should hold what it captured" defence does not survive §2
either: a stream of digits is not a different concept from the number it spells, so
holding both is one concept in two representations. And it defeated this module's
own consumers -- a context length nothing can compare with is a blob with a helpful
name.

The new witness asks both questions that were previously unaskable: that the
context length equals a number, and that utilization sits below unity. A regression
to text does not merely change a type, it stops that witness compiling.

Fifteen witnesses green.

Note for the record: review 59392 cites
w_the_served_id_advertises_a_quantisation_the_configured_path_contradicts, which
was renamed two commits earlier when that contradiction claim was narrowed to
non-derivability. That review read an older head; the finding it raised was still
live and is taken here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
…ed evidence

Side-chat ruling: take the historical fallback, and build it with the B5
authority split in one motion.

THE ACQUISITION CLAIM IS GONE. The previous shape called itself "one read
of one incarnation" and derived the refusal host and reported version from
it, which fixed two narrower defects and left the central one: the models
response, the metrics response and the evidence-route outcomes were still
independently assembled arguments. sole_constructor stops an external
record literal; it does not join values a caller supplied separately.
Nothing required one request batch, one endpoint, or one process, and
`observed_on: "2026-09-03"` cannot separate the several launches that came
and went on this pair that day.

The strong repair needs a retained per-request capture identity, and none
survived -- no attempt id, no per-route timestamps, no bodies. Minting one
now would convert a remembered procedure into fabricated provenance. So:
VllmHistoricalReadbackBundle over a VllmHistoricalFrontDoorScope, with
CrossRouteAcquisitionJoinUnestablished stamped by the constructor. That
type has NO established counterpart, so no later commit can flip this row
into a joined one, and no future rank or capacity observation can complete
it. A properly captured observation will be a different type through a
different constructor.

Both limitations are stamped and neither is computed. They answer
different questions -- could the observations be joined, and did anything
establish resolved configuration -- and collapsing them would lose the
first, which is what makes clear a later receipt cannot rehabilitate this
row.

THREE CALLER-AUTHORED POSITIVE PATHS REMOVED.
  - EvidenceRouteResolvedConfiguration carried a capture of three
    arbitrary strings from a public constructor. "garbage"/"0"/"anything"
    advanced the refusal stage because the ARM NAME said resolved. Gone;
    the arm belongs to the adapter that can produce a capture.
  - vllm_evidence_route(route:, kind:) let the kind be paired with the
    route by its caller, so (get_world_size, ResolvedParallelConfiguration)
    was writable and the stage law read that pairing. Replaced by a closed
    VllmParallelEvidenceRoute whose path, registration population and
    evidence kind are all projections.
  - The capacity decision echoed a caller-authored payload back as
    CapacityEvidenceAdmitted. It now classifies SOURCE KINDS -- authoritative
    versus insufficient -- and admits no evidence, which is honest for a
    module holding no capture.

The stage function is deleted rather than kept: with no configuration-
bearing response arm, it returns one value for every input a fixture can
write. A permanently-green check is worse than none, so the fact is stated
where it is known.

B5 SPLIT. extdeps.vllm.server scopes its laws against VllmSourceRevision
(the full upstream revision its source was read at). VllmSelfReportedVersion
moves downstream beside the readback that produced it -- an observation
this repository made is not a fact owned by the upstream. The two are not
joined: a version prefix is a lexical fact about a response body, never
runtime provenance.

Also corrected: the refusal said block_size is a minimum "across ranks",
conflating two axes. block_size is selected across resolved cache GROUPS;
num_blocks is reduced across WORKERS. A witness holds both spellings.

14 witnesses pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
@gunbai-bot

gunbai-bot Bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor Author

Rebuilt in one motion against the carried ruling (B2 fallback + B5 split), plus the three caller-mintable paths the rescore found.

B2 — the acquisition claim is gone. Verified the defect first: sole_constructor stops an external record literal but does not join values a caller supplied separately, and nothing required the models response, metrics response and evidence outcomes to come from one batch, endpoint or process. The strong repair needs a retained per-request capture identity; none survived this readback, so minting one would be fabricating provenance. VllmHistoricalReadbackBundle over a VllmHistoricalFrontDoorScope, with CrossRouteAcquisitionJoinUnestablished stamped by the constructor. That type has no established counterpart, so no later commit can flip the row and no future rank or capacity observation can complete it.

Both limitations are stamped, not collapsed — the join question and the resolved-configuration question are different, and only keeping both makes clear a later receipt cannot rehabilitate this bundle.

B3 — two caller-authored paths removed. EvidenceRouteResolvedConfiguration { capture } was three arbitrary strings from a public constructor; "garbage"/"0"/"anything" advanced the refusal stage because the arm name said resolved. Deleted. vllm_evidence_route(route:, kind:) let a caller write (get_world_size, ResolvedParallelConfiguration), which is exactly what the stage law read; replaced by a closed VllmParallelEvidenceRoute whose path, registration population and evidence kind are projections. Renamed to vllm_http_parallel_evidence_routes, since the host/Ray vantage is another possible source.

One judgement call worth flagging: I deleted the stage-derivation function rather than keeping the second arm reachable-in-principle. With no configuration-bearing response arm, it returns one value for every input a fixture can write — a permanently-green check, which this repo treats as worse than none because it gets cited as coverage. The coproduct keeps both arms as vocabulary; the fact is stated where it is known.

B4 — capacity classifies source kinds. decide_vllm_capacity_source_kind returns CapacitySourceAuthoritative | CapacitySourceInsufficient { cause }; no payload is admitted, because this module holds no capture. The group-versus-rank conflation is corrected — block_size is selected across resolved cache groups, num_blocks is reduced across workers — and a witness asserts both spellings so the wording cannot silently regress.

B5 — split. extdeps.vllm.server scopes its laws against VllmSourceRevision (1967a5627bc3710b680bbec24ecb99aaddedf22b); VllmSelfReportedVersion moved downstream beside the readback that produced it. Not joined — a version prefix is a lexical fact about a response body, never runtime provenance. Runtime-artifact and engine-incarnation identity are deliberately absent: they were not retained, and their absence is part of the bundle's ceiling.

B1 — body narrowed and synchronized. The alias section no longer says the id names a quantisation the deployment is not running or that an alias can lie; it claims what was observed, that the token is not derivable from the configured path and so the alias is not artifact identity. The Freshness section no longer says "one acquisition" or "another incarnation". Witness count corrected.

14 witnesses pass locally.

— sent from eager-pike-541

gunbc-ci-auto-heal and others added 3 commits September 3, 2026 21:00
…ts revision

Side-chat re-ruling on 5799acc. Two structural blockers plus the prose the
rebuild left behind.

THE VERSION EXISTED TWICE AND COULD DISAGREE. The scope derived a
reported_version from one VllmVersionRouteReadback and the bundle then
accepted another independently constructed one, so scope could say A while
the route said B with no construction preventing it -- one concept in two
independently authored representations. Worse, `/version` is a ROUTE-LOCAL
observation, and promoting what one route said into common scope for the
models route, the metrics gauge and the evidence observations asserts
exactly the cross-route join this bundle refuses to claim. The scope no
longer takes a version route and no longer carries a version;
VllmVersionRouteReadback carries VllmSelfReportedVersion directly.

THE CAPACITY RULING DID NOT CARRY ITS REVISION, while the route population
and the health fact both did -- so a consumer could hold "cache-config
fields are insufficient" with nothing saying which source that was read
from, and the body's claim that route, health AND capacity are each scoped
was not true. decide_vllm_capacity_source_kind now returns a
VllmCapacitySourceAssessment stamping established_against, and a witness
consumes the exact full revision from it, so a drift is red rather than a
silent re-scoping.

PROSE THAT STILL DESCRIBED THE DELETED TRANSACTION. The module header
still called this "a receipt about one acquisition"; a section was still
titled THE TRANSACTION IS WHY THESE CANNOT DISAGREE; and the smart-
constructor note still claimed a caller "cannot assemble a readback whose
parts came from different reads" -- which is precisely what
sole_constructor does NOT give, and the sentence was standing exactly
where the missing binding was. Each now says what the construction does:
the refusal host is derived from the scope, and that is the LIMIT of what
is bound.

The alias witness no longer describes a contradiction or an alias that can
lie. A served-model name may be arbitrary and a directory name does not
establish what its contents are quantised to, so the claim is
non-derivability, which is what was observed.

15 witnesses pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
Found by auditing this PR for the class a peer session just described:
an unproducible carrier reads as a finished model, so it never ranks for
inspection.

VllmHealthCoverageFact.backend was written once and read by nothing, and
MultiprocessingDistributedExecutor had no constructor and no consumer --
under an annotation claiming a consumer asking about another backend
"gets no answer here rather than a borrowed one". Prose cannot make that
true. The only reachable path to the coverage was the Ray row, so a
consumer holding a multiprocessing deployment would have read Ray's
answer straight off it, which is precisely what the sentence said could
not happen.

Coverage is now reached THROUGH the backend. vllm_health_coverage_for
returns HealthCoverageEstablished for Ray and
HealthCoverageNotEstablishedForBackend otherwise, so the second executor
arm is inhabited and the field is examined. The refusal is the content:
vLLM has more than one distributed executor, they do not share this
behaviour, and silence is the right answer for one whose source this
repository has not read.

This is the same defect class as a recorded-but-unexamined rank, and its
RED was UNAUTHORABLE before the lookup existed -- there was no question
to ask -- which is why three review rounds and two rulings passed over it.

16 witnesses pass. Main merged in; #10278 deleted the duplicate ledger
rows that were failing required-witnesses-build on every PR.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
@briansrls
briansrls merged commit 39fc622 into main Sep 3, 2026
5 of 6 checks passed
@briansrls
briansrls deleted the session/eager-pike-541-vllm-observed branch September 3, 2026 22:10
briansrls pushed a commit that referenced this pull request Sep 4, 2026
…d field nothing reads, under a sentence describing the wall it would build (#10315)

* File the class two lanes hit independently: prose asserting a wall nothing can produce

DESIGN section 4b requires a row per discovered class. This one earned it
by recurring five times across two unrelated lanes in one day.

THE CLASS: a carrier records a discriminating value, an annotation beside
it describes what that value prevents, and NO OPERATION ANYWHERE accepts
the value and can return a different answer because of it. The
declaration typechecks, the witnesses pass, and the sentence is not a lie
about intent -- it describes a wall that was never built. The field
corroborates the paragraph and the paragraph explains the field, so each
is the other's evidence and neither touches an executed path.

RECOGNITION RULE, about the question rather than the answer: name the
operation that TAKES this value and can return two results because of it.
If none exists, the field is inert and the prose is the only wall. The
mechanical tell is a field appearing exactly three times -- type
declaration, constructor parameter, constructor body -- and nowhere else.

WHY IT SURVIVES REVIEW, which is why it recurs: the discriminating RED is
UNAUTHORABLE. A reviewer hunting the missing negative case cannot write
one, because there is no call to make. Four instances passed three
independent review rounds and two side-chat rulings on #10264 and #10249
before any was found, and the fourth was found only after a peer session
described the shape from an unrelated lane.

RECEIPTS: the inert tensor_parallel_rank; a host vantage that cost a
wrapper call while its annotation claimed it had removed the free relabel;
a resolved-configuration capture of three arbitrary strings that advanced
a refusal stage on the strength of its ARM NAME; a health-coverage backend
written once and read by nothing under prose promising another backend
would get no answer. Plus a fifth from a peer lane on unrelated subject
matter -- "an annotation standing where a field should have been", beside
an advertised success surface with no construction anywhere.

BOUNDARIES DRAWN, because the nearest neighbour has a different remedy.
check_subject_shape_cannot_represent_the_state_the_check_detects has a
check that EXISTS and executes, unauthorable because the subject cannot
carry the state -- repair the subject's shape. Here no check exists at
all, unauthorable because there is nothing to call -- add the question.
And it is not specification-without-execution: there IS a consumer, green
by execution, so the general rule is satisfied BY the defect.

CEILING AND TRIGGER: both lanes found instances only after repairing
three or four smaller ones each, and both kept missing the next. The
trigger is a lens over the Node tree asking whether any function in the
closure destructures each declared field -- the same read the namespace
authority performs, so it costs no substrate edit. Until then the class
is mitigatable by review diligence, which both lanes have measured to be
insufficient.

Projection regenerated via the generated-artifact gate; 5 agreement
witnesses pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G

* State the rule as the census can decide it, and refuse the second borrowed receipt

Side-chat review found that this row's rule and this row's instrument were
not asking the same question, and that a second receipt did not belong.

THE RULE AND THE LENS DISAGREED.

The recognition rule asked whether an operation "can return two different
results because of" the field. That is semantic INFLUENCE. The next-rung
lens asked only whether some function destructures the field. That is
SYNTACTIC USE. They are not equivalent: a function may bind a field into a
dead binding, or compare it inside an arm already decided, leaving the
answer invariant while a use census reports the wall present.

Shipping both would have reproduced, inside this row's own enforcement,
exactly the false coverage the row exists to name.

Narrowed to the mechanically decidable population: no RESOLVED,
NON-CONSTRUCTION READ anywhere in the field's dependent production closure.
Both qualifiers carry weight. RESOLVED, because a consumer may sit in
another module or reach the field through an alias, so a textual miss would
report inertness that is not there. NON-CONSTRUCTION, because the
constructor body reads every field by definition -- a census counting that
read finds no inert field anywhere and is permanently green.

The three-appearance tell is demoted to what it is: a candidate generator,
the cheapest way to find one, never the adjudicator. The influence residue
goes to live_argument_threaded_past_the_arm_that_decides, which already
exists, rather than being absorbed here.

THE OOBE RECEIPT IS A DIFFERENT CLASS.

`OobeBrowserSessionReady` was carried as a fifth receipt. Whole-tree
resolution found NO construction anywhere, so no field instance was ever
produced and then ignored. That contradicts this row's own boundary
sentence, which requires a consumer that EXISTS, executes on every run, and
is green while one field it carries goes uninterrogated. The remedies
diverge too: an unread field needs a question that consumes it, an
unproducible success surface needs a construction path or the deletion of
the advertised state.

Two candidates from two lanes are now admitted by resemblance and refused by
the rule. Both refusals stay in the row. That pair is the evidence the
boundary is enforced by applying the rule rather than by asserting it.

THE CAPTURE RECEIPT'S GRAIN WAS TOO COARSE.

The outer response discriminant WAS read, and reading it advanced the stage.
What nothing examined were the three strings inside the capture payload. The
receipt now says payload rather than coproduct, because a receipt that
misplaces its own subject is not a receipt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G

---------

Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant