Repository navigation
Read the vLLM front door back, and refuse to call it a TP realization - #10249
Conversation
Step 2 of making serving_converge_slice_wet safe: describe what is actually on the withheld pair. The load-bearing half is what the readback is NOT sufficient to claim. WHAT WAS OBSERVED, first-hand over HTTP rather than by report: /version gives a vLLM build string, /v1/models gives owned_by vllm with the official DeepSeek-V4-Flash snapshot as the model root, and the vllm:cache_config_info gauge on /metrics gives the KV dtype, memory utilization and prefix-caching flag. THE SERVED ID LIES AND THE ARTIFACT ROOT SETTLES IT. The id reads "deepseek-v4-flash:iq3s-split", naming a quantisation this deployment is not running -- it is a --served-model-name alias chosen so the router's model name matched the other pair. Two sessions concluded the wrong thing about the weights from that id, once in a design ruling that then propagated. The witness pins the disagreement itself rather than the corrected value, so a later edit that tidies the id to match the root cannot quietly erase the evidence that an alias can lie. WHY THERE IS NO TENSOR-PARALLEL REALIZATION HERE, AND WHY THE REFUSAL IS WORDED AS IT IS. It would be easy to refuse because the second host did not answer on the front-door port. That models the absence of something never promised: a tensor-parallel worker is not required to expose its own OpenAI-compatible server, one front door is the normal shape, and such a refusal would go green the moment an unrelated server appeared on that port. The missing thing is a POSITIVE joined-rank population receipt. The refusal is a two-arm coproduct because the stages are different questions -- you cannot ask whether every expected rank is present until resolved parallel configuration tells you how many are expected, and that expected set must be derived rather than supplied by the caller, or the caller authors both sides of its own join. This observation never reached the configuration stage, so it refuses at the earlier arm. The routes that could answer it (/server_info, /get_world_size, /collective_rpc) all 404'd, which under the new extdeps route split means the server was launched without development mode. That is NOT a reason to relaunch it: /collective_rpc executes an arbitrary named method across every worker and upstream documents development mode as unfit for production. Spending a real safety property to buy a modelling convenience is the wrong trade, and the withholding this feeds holds without it. CAPACITY IS ABSENT RATHER THAN ESTIMATED, and this is a correction. The gauge also exposed num_gpu_blocks and a resolved minimum block_size; I multiplied them, called the product the live token pool, and published it. It is not one -- the gauge reports CacheConfig fields, while a multi-group cache layout has the engine compute capacity group-aware and log it at startup after reducing every worker to the minimum block count across ranks. The two numbers are real observations and are omitted anyway, because a block count sitting in a capacity carrier will be multiplied again by someone with less context. vLLM gets its own extdeps module rather than a variant in a serving-runtime enum, which is what extdeps.ollama.api's own note asks for from the other side: adding vLLM "must not require editing this file". It did not. EXECUTED EVIDENCE. Six witnesses green, then two mutations on the production path: asserting the later refusal stage, and putting an unanswered port into the refusal wording. Exactly the two targeted claims went red and the other four held, so each mutation broke its own subject rather than the module. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
…ity law
Eight findings, all real, several of them my recurring overclaim pattern.
STRUCTURAL BINDING. The readback and the refusal were separate rows that each
named a host, so one could be edited to srv5 while the other said srv6 and every
witness stayed green. They are now one VllmReadbackTransaction whose constructor
DERIVES the refusal from the acquisition, so a divergent pair has no constructor.
Construction over validation, rather than a witness checking two rows agree.
THE STAGE IS DERIVED, NOT ASSERTED. Which arm of the missing-evidence coproduct
applies now follows from whether any parallel-evidence route answered, so the
earlier arm is a consequence rather than a claim.
THE CAPACITY AUTHORITY WAS MALFORMED. It was a coproduct whose arms were
`GroupAwareStartupLogLine` and `BlockCountTimesMinimumBlockSizeIsNotCapacity`.
Those are not alternatives -- both hold at once -- so selecting one left the
rejection carried by prose and by the spelling of an uninhabited arm. It is now a
decision over proposed evidence with both directions executable. The refusal also
states the GENERAL law rather than an arithmetic complaint: for a uniform layout
the product may coincide with the right answer, so the true rule is that those two
cache-config fields are not a generally SUFFICIENT authority.
Removed `CapacityFromGroupAwareStartupLog { pool_tokens }`, which was a future
positive arm with no parser and no sealed constructor -- inconsistent with my own
stated reason for not landing the joined-rank receipt nouns. Removed
`vllm_observed_capacity_authority()`, which restated the extdeps constant under a
second name.
THREE OVERCLAIMS RETRACTED. (1) `model_artifact_root` is renamed
`configured_model_path`: /v1/models reports the path the server was CONFIGURED
with, which is evidence about intent and none about bytes -- nothing here digests
the weights. I had published "the weights really are the official checkpoint,
confirmed first-hand" upward, and that was not established. (2) The pull request
attributed max_model_len to the cache-config gauge; it came from /v1/models. Field
provenance is now structural -- records are grouped by source route, so a field
cannot be described as coming from a route it does not sit under. (3) The module
said three 404s mean "launched without development mode". A 404 does not read a
launch flag. It establishes the evidence was unreachable, which is all the refusal
needs, so the refusal now rests on that and survives the inference being wrong.
VERSION AND BACKEND BINDING. A repository-root URI anchors the subject but does
not ground version-sensitive facts, so route population, health coverage and the
capacity law now carry the build they were established against. The health fact is
additionally backend-bound: it is a property of one executor's engine-health
implementation, not of the route, so a consumer asking about another backend gets
no answer rather than a borrowed one.
`AlwaysServed` became `OrdinarilyRegistered`. "Always" claims every such route
answers on every deployment, which one probed server cannot establish -- /metrics
is suppressible by launch configuration.
WITNESSES NOW CONSUME DECISIONS RATHER THAN READ ROWS. A type name, a coproduct
arm and a record field are type dependencies that consume nothing, and the previous
suite destructured data declarations, so each of these stayed green: flipping the
Ray health-coverage row, flipping the capacity authority, and pointing the
observation at srv5 while the refusal said srv6. Nine of ten claims now call a
function, with hermetic fixtures driving the constructor in both directions.
EXECUTED EVIDENCE. Ten green, then the three mutations the reviewer named as
previously undetectable:
FAIL w_the_refusal_host_is_derived_from_the_acquisition_it_belongs_to
FAIL w_the_group_aware_startup_capture_is_the_admitted_capacity_evidence
FAIL w_ray_health_does_not_probe_distributed_workers
Each caught by its own claim, the other seven holding.
NOT DONE HERE, DELIBERATELY: the vLLM rank-population decision kernel the reviewer
ruled startable now. It is a separate construction with its own hermetic fixture
matrix, and the sequencing it gave puts the readback repair first.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
…swer A semantic bug in the derivation I added in the previous commit, and it errs in the direction that overstates what is known. `vllm_missing_evidence_for` advanced to the population question whenever ANY parallel-evidence route answered. The three routes do not establish the same thing. /get_world_size returns a PRODUCT -- a world size of 2 is consistent with TP=2, PP=2 and DP=2 alike, so on its own it narrows the space without deciding it. /collective_rpc fans a named method across workers, which is evidence about population and says nothing about how the parallelism decomposes. Only /server_info returns the resolved configuration. So either of the other two could answer while resolved parallel configuration stayed entirely unobserved, and the refusal would then claim the configuration was in hand and name the LATER missing fact -- asserting knowledge nobody had, in a module whose entire subject is not doing that. THE KIND IS NOW CARRIED ON THE ROUTE RATHER THAN DERIVED FROM ITS PATH. The first fix dispatched on the path literal, which makes a semantic fact depend on a spelling: rename the route upstream and the kind silently becomes the fallback arm with nothing red. VllmEvidenceRoute pairs route and kind at construction, so a route cannot enter the evidence population without its kind being stated. Three witnesses, and the third is what stops the first two from being satisfiable by a degenerate rule: a world-size-plus-fan-out answer must NOT advance the stage; a configuration answer alone MUST advance it (otherwise "never advance" would pass the first while making the later arm unreachable); and the kinds must actually differ (otherwise both pass if every route reports the configuration-bearing kind). Thirteen witnesses green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
…mixed build MIXED-BUILD TRANSACTIONS WERE CONSTRUCTIBLE. The acquisition carried a version and the version route carried one, so a transaction could name two different builds while looking like a single read -- the exact cross-incarnation reuse this module exists to prevent, reachable inside the record meant to prevent it. The constructor now takes the version route and DERIVES the acquisition's version from it, so disagreement has no constructor. `answered: Bool` CONFLATED THREE STATES AND HID THE DANGEROUS ONE. A 404, a 200 nobody parsed, and a parsed resolved configuration are not one bit. The middle case is where "we did not look" becomes "we know": a body that was never read is not knowledge, and a Bool recorded it identically to an answer. The outcome now carries a response coproduct, and only the parsed-configuration arm moves the refusal stage. That makes the later stage unreachable from production today -- correct, as no parser for that payload exists, which is precisely what the Bool let me paper over. THE CAPACITY ADMISSION WAS VACUOUS. `GroupAwareStartupLogCapture` was a bare arm with no payload, so admitting it asserted that a KIND of evidence would be acceptable -- something no capacity claim can be built from. The positive direction was permanently unreachable and the decision was a refusal wearing two arms. The proposal and the admission now carry the parsed capture, so admitting requires having one. THE VERSION IS SELF-REPORTED AND THE TYPE NOW SAYS SO. `VllmBuildIdentity` claimed more than /version gives: a process reports what it says about itself, nothing digests the executed image, and nothing cross-checks it. Renamed `VllmSelfReportedVersion` -- adequate for distinguishing incarnations and scoping version-sensitive facts, inadequate for any claim about what code is running. AND THE ALIAS CLAIM IS NARROWED TO WHAT A SUBSTRING TEST ESTABLISHES. I asserted the served id and configured path CONTRADICT each other. They do not: a path need not name a quantisation at all, so its absence is not evidence the id's token is false. What is established is weaker and sufficient -- the id is not derivable from the path, so it must not be read as an artifact identity. The refusal of the id as an artifact claim stands; the contradiction does not. Fourteen witnesses green, including the case a Bool could not express: an unparsed 200 on the configuration-bearing route leaves the refusal exactly where a 404 does. ONE FINDING NOT TAKEN, AND FLAGGED RATHER THAN SILENTLY DROPPED: typing the numeric readbacks (max_model_len, gpu_memory_utilization) as std.measure carriers. The two reviewers disagree here -- the other holds that these are verbatim front-door capture text at the extdeps boundary and that converting them inward is what the raw-scalar rule exists to prevent. I have left them as captured strings and am carrying the disagreement to the operator rather than picking a side by edit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
Both reviewers now agree on this, having previously disagreed. The GitHub reviewer had approved these as verbatim extdeps-boundary captures and has reversed: string-as-number is the same class as raw-scalar "one step worse", because it loses ordering and arithmetic on top of the unit. The side chat listed them as untyped throughout. I had flagged the disagreement rather than picking a side by edit; it is now resolved, so I am taking it. max_model_len is a TokenCount and gpu_memory_utilization a BasisPoint, both against std.measure's existing authorities rather than fresh ones. Basis points rather than Percent because Percent would round a two-decimal fraction into a Nat and lose the distinction between 0.86 and 0.855, which vLLM accepts as different. THE ABSENT PARSER IS WHAT DECIDES THIS, not a preference between two defensible shapes. The tempting middle -- carry the wire text AND a typed value -- is worse than either end, because nothing in this repository parses one into the other. The two fields would be independently authored and free to drift, which is exactly the validation shape this module has spent four rounds replacing with construction. So the measure is the fact; when a parser exists it can produce one FROM a capture rather than beside it. The "capture module should hold what it captured" defence does not survive §2 either: a stream of digits is not a different concept from the number it spells, so holding both is one concept in two representations. And it defeated this module's own consumers -- a context length nothing can compare with is a blob with a helpful name. The new witness asks both questions that were previously unaskable: that the context length equals a number, and that utilization sits below unity. A regression to text does not merely change a type, it stops that witness compiling. Fifteen witnesses green. Note for the record: review 59392 cites w_the_served_id_advertises_a_quantisation_the_configured_path_contradicts, which was renamed two commits earlier when that contradiction claim was narrowed to non-derivability. That review read an older head; the finding it raised was still live and is taken here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
…ed evidence
Side-chat ruling: take the historical fallback, and build it with the B5
authority split in one motion.
THE ACQUISITION CLAIM IS GONE. The previous shape called itself "one read
of one incarnation" and derived the refusal host and reported version from
it, which fixed two narrower defects and left the central one: the models
response, the metrics response and the evidence-route outcomes were still
independently assembled arguments. sole_constructor stops an external
record literal; it does not join values a caller supplied separately.
Nothing required one request batch, one endpoint, or one process, and
`observed_on: "2026-09-03"` cannot separate the several launches that came
and went on this pair that day.
The strong repair needs a retained per-request capture identity, and none
survived -- no attempt id, no per-route timestamps, no bodies. Minting one
now would convert a remembered procedure into fabricated provenance. So:
VllmHistoricalReadbackBundle over a VllmHistoricalFrontDoorScope, with
CrossRouteAcquisitionJoinUnestablished stamped by the constructor. That
type has NO established counterpart, so no later commit can flip this row
into a joined one, and no future rank or capacity observation can complete
it. A properly captured observation will be a different type through a
different constructor.
Both limitations are stamped and neither is computed. They answer
different questions -- could the observations be joined, and did anything
establish resolved configuration -- and collapsing them would lose the
first, which is what makes clear a later receipt cannot rehabilitate this
row.
THREE CALLER-AUTHORED POSITIVE PATHS REMOVED.
- EvidenceRouteResolvedConfiguration carried a capture of three
arbitrary strings from a public constructor. "garbage"/"0"/"anything"
advanced the refusal stage because the ARM NAME said resolved. Gone;
the arm belongs to the adapter that can produce a capture.
- vllm_evidence_route(route:, kind:) let the kind be paired with the
route by its caller, so (get_world_size, ResolvedParallelConfiguration)
was writable and the stage law read that pairing. Replaced by a closed
VllmParallelEvidenceRoute whose path, registration population and
evidence kind are all projections.
- The capacity decision echoed a caller-authored payload back as
CapacityEvidenceAdmitted. It now classifies SOURCE KINDS -- authoritative
versus insufficient -- and admits no evidence, which is honest for a
module holding no capture.
The stage function is deleted rather than kept: with no configuration-
bearing response arm, it returns one value for every input a fixture can
write. A permanently-green check is worse than none, so the fact is stated
where it is known.
B5 SPLIT. extdeps.vllm.server scopes its laws against VllmSourceRevision
(the full upstream revision its source was read at). VllmSelfReportedVersion
moves downstream beside the readback that produced it -- an observation
this repository made is not a fact owned by the upstream. The two are not
joined: a version prefix is a lexical fact about a response body, never
runtime provenance.
Also corrected: the refusal said block_size is a minimum "across ranks",
conflating two axes. block_size is selected across resolved cache GROUPS;
num_blocks is reduced across WORKERS. A witness holds both spellings.
14 witnesses pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
|
Rebuilt in one motion against the carried ruling (B2 fallback + B5 split), plus the three caller-mintable paths the rescore found. B2 — the acquisition claim is gone. Verified the defect first: Both limitations are stamped, not collapsed — the join question and the resolved-configuration question are different, and only keeping both makes clear a later receipt cannot rehabilitate this bundle. B3 — two caller-authored paths removed. One judgement call worth flagging: I deleted the stage-derivation function rather than keeping the second arm reachable-in-principle. With no configuration-bearing response arm, it returns one value for every input a fixture can write — a permanently-green check, which this repo treats as worse than none because it gets cited as coverage. The coproduct keeps both arms as vocabulary; the fact is stated where it is known. B4 — capacity classifies source kinds. B5 — split. B1 — body narrowed and synchronized. The alias section no longer says the id names a quantisation the deployment is not running or that an alias can lie; it claims what was observed, that the token is not derivable from the configured path and so the alias is not artifact identity. The Freshness section no longer says "one acquisition" or "another incarnation". Witness count corrected. 14 witnesses pass locally. — sent from eager-pike-541 |
…ts revision Side-chat re-ruling on 5799acc. Two structural blockers plus the prose the rebuild left behind. THE VERSION EXISTED TWICE AND COULD DISAGREE. The scope derived a reported_version from one VllmVersionRouteReadback and the bundle then accepted another independently constructed one, so scope could say A while the route said B with no construction preventing it -- one concept in two independently authored representations. Worse, `/version` is a ROUTE-LOCAL observation, and promoting what one route said into common scope for the models route, the metrics gauge and the evidence observations asserts exactly the cross-route join this bundle refuses to claim. The scope no longer takes a version route and no longer carries a version; VllmVersionRouteReadback carries VllmSelfReportedVersion directly. THE CAPACITY RULING DID NOT CARRY ITS REVISION, while the route population and the health fact both did -- so a consumer could hold "cache-config fields are insufficient" with nothing saying which source that was read from, and the body's claim that route, health AND capacity are each scoped was not true. decide_vllm_capacity_source_kind now returns a VllmCapacitySourceAssessment stamping established_against, and a witness consumes the exact full revision from it, so a drift is red rather than a silent re-scoping. PROSE THAT STILL DESCRIBED THE DELETED TRANSACTION. The module header still called this "a receipt about one acquisition"; a section was still titled THE TRANSACTION IS WHY THESE CANNOT DISAGREE; and the smart- constructor note still claimed a caller "cannot assemble a readback whose parts came from different reads" -- which is precisely what sole_constructor does NOT give, and the sentence was standing exactly where the missing binding was. Each now says what the construction does: the refusal host is derived from the scope, and that is the LIMIT of what is bound. The alias witness no longer describes a contradiction or an alias that can lie. A served-model name may be arbitrary and a directory name does not establish what its contents are quantised to, so the claim is non-derivability, which is what was observed. 15 witnesses pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
Found by auditing this PR for the class a peer session just described: an unproducible carrier reads as a finished model, so it never ranks for inspection. VllmHealthCoverageFact.backend was written once and read by nothing, and MultiprocessingDistributedExecutor had no constructor and no consumer -- under an annotation claiming a consumer asking about another backend "gets no answer here rather than a borrowed one". Prose cannot make that true. The only reachable path to the coverage was the Ray row, so a consumer holding a multiprocessing deployment would have read Ray's answer straight off it, which is precisely what the sentence said could not happen. Coverage is now reached THROUGH the backend. vllm_health_coverage_for returns HealthCoverageEstablished for Ray and HealthCoverageNotEstablishedForBackend otherwise, so the second executor arm is inhabited and the field is examined. The refusal is the content: vLLM has more than one distributed executor, they do not share this behaviour, and silence is the right answer for one whose source this repository has not read. This is the same defect class as a recorded-but-unexamined rank, and its RED was UNAUTHORABLE before the lookup existed -- there was no question to ask -- which is why three review rounds and two rulings passed over it. 16 witnesses pass. Main merged in; #10278 deleted the duplicate ledger rows that were failing required-witnesses-build on every PR. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
…d field nothing reads, under a sentence describing the wall it would build (#10315) * File the class two lanes hit independently: prose asserting a wall nothing can produce DESIGN section 4b requires a row per discovered class. This one earned it by recurring five times across two unrelated lanes in one day. THE CLASS: a carrier records a discriminating value, an annotation beside it describes what that value prevents, and NO OPERATION ANYWHERE accepts the value and can return a different answer because of it. The declaration typechecks, the witnesses pass, and the sentence is not a lie about intent -- it describes a wall that was never built. The field corroborates the paragraph and the paragraph explains the field, so each is the other's evidence and neither touches an executed path. RECOGNITION RULE, about the question rather than the answer: name the operation that TAKES this value and can return two results because of it. If none exists, the field is inert and the prose is the only wall. The mechanical tell is a field appearing exactly three times -- type declaration, constructor parameter, constructor body -- and nowhere else. WHY IT SURVIVES REVIEW, which is why it recurs: the discriminating RED is UNAUTHORABLE. A reviewer hunting the missing negative case cannot write one, because there is no call to make. Four instances passed three independent review rounds and two side-chat rulings on #10264 and #10249 before any was found, and the fourth was found only after a peer session described the shape from an unrelated lane. RECEIPTS: the inert tensor_parallel_rank; a host vantage that cost a wrapper call while its annotation claimed it had removed the free relabel; a resolved-configuration capture of three arbitrary strings that advanced a refusal stage on the strength of its ARM NAME; a health-coverage backend written once and read by nothing under prose promising another backend would get no answer. Plus a fifth from a peer lane on unrelated subject matter -- "an annotation standing where a field should have been", beside an advertised success surface with no construction anywhere. BOUNDARIES DRAWN, because the nearest neighbour has a different remedy. check_subject_shape_cannot_represent_the_state_the_check_detects has a check that EXISTS and executes, unauthorable because the subject cannot carry the state -- repair the subject's shape. Here no check exists at all, unauthorable because there is nothing to call -- add the question. And it is not specification-without-execution: there IS a consumer, green by execution, so the general rule is satisfied BY the defect. CEILING AND TRIGGER: both lanes found instances only after repairing three or four smaller ones each, and both kept missing the next. The trigger is a lens over the Node tree asking whether any function in the closure destructures each declared field -- the same read the namespace authority performs, so it costs no substrate edit. Until then the class is mitigatable by review diligence, which both lanes have measured to be insufficient. Projection regenerated via the generated-artifact gate; 5 agreement witnesses pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * State the rule as the census can decide it, and refuse the second borrowed receipt Side-chat review found that this row's rule and this row's instrument were not asking the same question, and that a second receipt did not belong. THE RULE AND THE LENS DISAGREED. The recognition rule asked whether an operation "can return two different results because of" the field. That is semantic INFLUENCE. The next-rung lens asked only whether some function destructures the field. That is SYNTACTIC USE. They are not equivalent: a function may bind a field into a dead binding, or compare it inside an arm already decided, leaving the answer invariant while a use census reports the wall present. Shipping both would have reproduced, inside this row's own enforcement, exactly the false coverage the row exists to name. Narrowed to the mechanically decidable population: no RESOLVED, NON-CONSTRUCTION READ anywhere in the field's dependent production closure. Both qualifiers carry weight. RESOLVED, because a consumer may sit in another module or reach the field through an alias, so a textual miss would report inertness that is not there. NON-CONSTRUCTION, because the constructor body reads every field by definition -- a census counting that read finds no inert field anywhere and is permanently green. The three-appearance tell is demoted to what it is: a candidate generator, the cheapest way to find one, never the adjudicator. The influence residue goes to live_argument_threaded_past_the_arm_that_decides, which already exists, rather than being absorbed here. THE OOBE RECEIPT IS A DIFFERENT CLASS. `OobeBrowserSessionReady` was carried as a fifth receipt. Whole-tree resolution found NO construction anywhere, so no field instance was ever produced and then ignored. That contradicts this row's own boundary sentence, which requires a consumer that EXISTS, executes on every run, and is green while one field it carries goes uninterrogated. The remedies diverge too: an unread field needs a question that consumes it, an unproducible success surface needs a construction path or the deletion of the advertised state. Two candidates from two lanes are now admitted by resemblance and refused by the rule. Both refusals stay in the row. That pair is the evidence the boundary is enforced by applying the rule rather than by asserting it. THE CAPTURE RECEIPT'S GRAIN WAS TOO COARSE. The outer response discriminant WAS read, and reading it advanced the stage. What nothing examined were the three strings inside the capture payload. The receipt now says payload rather than coproduct, because a receipt that misplaces its own subject is not a receipt. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G --------- Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Step 2 of making
serving_converge_slice_wetsafe: describe what is actually on the pair #10225 withholds. The load-bearing half is what the readback is not sufficient to claim.What was observed
First-hand over HTTP, not by report:
/version/v1/modelsowned_by: vllm, a configured model path,max_model_len262144/metrics(vllm:cache_config_info)cache_dtypefp8_ds_mla,gpu_memory_utilization0.86, prefix caching onRecords are grouped by the route they came from, so provenance is structural: a field cannot be described as coming from a route it does not sit under.
The served alias and the configured path are two separate facts. The id reads
deepseek-v4-flash:iq3s-splitand carries a quantisation token that is not derivable from the configured path — a--served-model-namealias chosen so the router's model name matched the other pair. That is enough to refuse the id as artifact identity. It is not enough to conclude the alias is false: a served-model name is allowed to be arbitrary, and a directory name does not establish what its contents are quantised to. Two sessions concluded the wrong thing about the weights from that id, once in a design ruling that propagated; the witness pins the non-derivability, which is the part that was observed.Why there is no TP realization here
Refusing because the second host did not answer on the front-door port would model the absence of something never promised: a tensor-parallel worker is not required to expose its own OpenAI-compatible server, and such a refusal would go green the moment an unrelated server appeared on that port.
What is missing is a positive joined-rank population receipt. The refusal is two-stage because the questions differ: you cannot ask whether every expected rank is present until resolved parallel configuration says how many are expected, and that expected set must be derived, not supplied by the caller.
The route identity decides its kind; neither is supplied by a caller.
/get_world_sizereturns a product consistent with TP=2, PP=2 and DP=2 alike;/collective_rpcreports population without saying how parallelism decomposes. Only/server_inforeturns resolved configuration. An earlier version paired route and kind at construction, which left the relation caller-authored —(get_world_size, ResolvedParallelConfiguration)was writable, and the stage law read that pairing. It is now a projection of a closedVllmParallelEvidenceRoute, so no combination can be authored.There is no parsed-configuration arm, and its absence is the point. The previous
EvidenceRouteResolvedConfiguration { capture }was built from three arbitrary strings by a public constructor: a caller could write"garbage","0","anything", wrap it, and advance the refusal stage because the arm name said the configuration had been resolved. That arm belongs to the adapter that can actually produce a capture.All three evidence routes 404'd. That is not a reason to relaunch the server.
/collective_rpcexecutes an arbitrary named method across every worker and upstream documents development mode as unfit for production. The withholding this feeds holds without it.Retractions
/v1/modelsreports the path the server was configured with — evidence about intent, none about bytes. Nothing here digests the weights. The field is nowconfigured_model_path.max_model_lenattributed to the cache-config gauge. It came from/v1/models.num_gpu_blocks × block_size). Not this build's capacity law, and the two factors are not even on the same axis:block_sizeis selected as the minimum across resolved cache groups whilenum_blocksis separately reduced across workers, so the product joins values from different axes — and a multi-group layout has the engine compute capacity group-aware and log it at startup. The bundle carries no capacity field at all: a block count in a capacity carrier gets multiplied again by the next reader.Structure
sole_constructorstops an external record literal; it does not join values a caller supplied separately, andobserved_on: "2026-09-03"cannot separate the several launches that came and went on this pair that day.CrossRouteAcquisitionJoinUnestablishedis stamped by the constructor and has no established counterpart — no later commit can flip this row into a joined one, and no future rank or capacity observation can complete it. A properly captured observation will be a different type through a different constructor.CapacityEvidenceAdmitted. This module holds no capture, so it says which kind of source would be authoritative and stops there.extdeps.vllm.serverscopes its route, health and capacity laws against the full revision its source was read at;VllmSelfReportedVersionmoved downstream beside the readback that produced it, because an observation this repository made is not a fact owned by the upstream. The two are not joined — a version prefix is a lexical fact about a response body, never runtime provenance.RayDistributedExecutor, whose engine-health implementation returns without contacting workers.OrdinarilyRegistered, notAlwaysServed— one probed server cannot establish that a route answers on every deployment.Executed evidence
Fourteen witnesses green, driving constructors and decisions with hermetic fixtures rather than destructuring the pinned row. Both capacity directions execute; the group-versus-rank axes are asserted separately; the stage-derivation function was deleted rather than kept, because with no configuration-bearing response arm it returns one value for every input a fixture can write, and a permanently-green check is worse than none.
Freshness
The front door answered on 2026-09-03 and had stopped answering within the hour;
.225/.226are now silent while production.232/.233are up. This is a bundle of route-local observations associated with one front-door investigation — not one acquisition and not one identified incarnation, neither of which the retained evidence establishes. It supports what each route reported. It does not support that all values belonged to one engine process, that all requests preceded any restart, or that the self-reported version identifies executed bytes.🤖 Generated with Claude Code
https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G