Repository navigation
Sparks Pair (P0->P5) - #9897
Sparks Pair (P0->P5)#9897
Conversation
…ot be narrowed by a distributor
A probe of one distributor's registry returned 404 and was reported as "no local
weights are published". That inference excluded three openly-published releases
from the candidate population -- including the two highest-scoring open models
available -- and nothing downstream could recover them, because excluding a
candidate from the search is not visible as a ranking error.
The defect was structural, not a lapse. `ModelArtifactProvenance` spells stock
provenance as `PulledFromUpstream { publisher, reference: OllamaModelRef }`, so
one distributor's vocabulary was the only way to say where a model came from and
that distributor's silence had nowhere to land except as absence of the model.
The same file already argued the principle and applied it correctly to the tuning
arm, which carries a `ContentHash` rather than a reference precisely because "a
lineage built on moving pointers records a story rather than a fact".
gunbc.model.publication separates the four facts that were fused: release,
weight publication, distribution channel, and packaging. Open-weight-ness is a
CONSTRUCTOR of `ModelRelease`, so it is answered by matching the release and no
channel observation participates -- there is no expressible path from a 404 to a
change in that answer, rather than a check that would catch one. `ChannelPresence`
has no arm spelling unqualified absence: `AbsentFromChannel` names the channel it
is absent from, so "this model is unavailable" is not a sentence the type can
produce.
gunbc.model.population makes the narrowing invariant structural. A later stage is
constructed from its parent by `filter`, so it cannot introduce a member the
parent lacked, and refusing a stage cannot reach back into the parent because the
parent is a separate immutable value the narrowing consumed. There is no writable
state for the regression to be written into.
`ReleaseDiscoverySource` carries a declared `CoverageScope`, and
`open_weight_completeness` REFUSES when every source is a single distributor or an
operator roster. The layer split alone does not repair this: the original error was
a discovery error, and a census that only ever probes one distributor stays
distributor-shaped while reporting that shape as completeness.
Parameter counts are three sourced fields rather than one reconciled number. The
two circulating totals for the fixture release differ by exactly the size of a
separately-shipped speculative draft, so which one is meant decides whether a
quantization fits a host.
Evidence, established by execution: `w_all_population_claims_hold` returns true
(eight claims -- four forbidden edges with the release that actually broke as the
fixture, the discovery-coverage refusal, and two positive controls so the
refusals discriminate). `w_red_arm_forbidden_edge_must_evaluate_false` returns
false, asserting the forbidden edge directly; without it the suite would be a
conjunction of things that happen to be true.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HL72ndT9xg2dZEb6J8a2Gr
…trol that makes it a selector choice.dag landed with no executed consumer, which is the specification-without- execution trap: it parsed, so the module index accepted it, and nothing ran it. The fixture is measured rather than plausible. Capacity is the kernel-visible MemTotal read from the running nodes (121.69 GiB), beside the vendor nominal (128 GiB) so the ~6.31 GiB firmware carveout is visible rather than rediscovered; quoting the nominal figure is what made a 128.1 GB build look feasible. KV is exact rather than estimated: the served model reports 43 blocks, ONE latent KV head, and key/value lengths of 512, so a token costs 88,064 bytes. The two candidates are the builds actually installed, at their real byte sizes. The claims are the operating conclusions, mechanised. At a 400k context floor the HIGHER-quality build is rejected on memory and the selector returns the 2-bit one -- precision and the context floor compete for one node's memory, and the function says so from the numbers instead of from an argument. The load-bearing test is the flip control: with the context floor dropped, the selector must choose the HIGHER-quality build. Without it every other claim is satisfied by a function that returns its first argument, and the suite would establish nothing about whether quality is maximized at all. The refusal arm is exercised too: with an unreachable floor the selector returns NoCandidateAdmissible carrying both rejections, so the binding axis is read rather than argued. A best-effort pick there would reintroduce exactly the silent degradation this module exists to remove. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HL72ndT9xg2dZEb6J8a2Gr
…orpus-wide variant names Two defects, both found by machinery rather than by reading. REVIEW FINDING (review 58079): every unit-bearing field in choice.dag was a bare Int -- capacities, weights, the KV rate, prefill throughput -- instead of consuming std/measure.dag carriers. Fixed at every site the review named plus the two it flagged as adjacent smells, since they are the same class: ByteSize for capacity/weights/KV, TokenCount for context, and a rate carrier for prefill. The verdict surface is converted too, so DoesNotFitMemory and the admissible arm no longer propagate flat scalars outward. quality_rank stays Int deliberately: it is a dimensionless ordinal, not a unit quantity. TokensPerSecond did not exist, so it is added to std.measure rather than minted locally -- consuming the single authority is the whole point of the finding. Its annotation records why it is a rate and not a TokenCount, and warns that prefill and decode rates share the type so the FIELD NAME has to separate them. The finding is sharper than a style rule and it bit this lane repeatedly: the session it came from produced four distinct wrong answers by attaching a correct number to the wrong population, including quoting a vendor-nominal memory figure as allocatable capacity, which made an over-capacity build look feasible. ByteSize and TokenCount make one of those classes unwritable instead of something a reader has to keep catching. CI FAILURE, same push: the required floor lane failed while the build lane passed. Cause was `Admissible` and `Rejected` as CandidateVerdict arms. Variant names resolve corpus-wide, so those two shadowed arms in four unrelated modules -- extdeps.tools.jq, extdeps.languages.markdown, extdeps.bmc.openbmc_fan_control and extdeps.provisioning.ubuntu_install_media_fetch -- none of which this change otherwise touches. That is why only the whole-corpus lane caught it, twenty minutes in. Renamed to ServingCandidateAdmissible / ServingCandidateRejected; CandidateRejected was rejected as a replacement because it collides in turn with the self-host door-observation modules. Verified by an exhaustive scan of all 38 variant arms introduced by these modules against the rest of the corpus, which now reports zero collisions. A local canary compile cannot establish this: compiling a victim module scopes to its own closure and never loads these modules, so only a whole-corpus pass observes the clash. Evidence after both fixes: selector aggregate true, selector flip control true (the control that distinguishes a quality-maximizing selector from one returning its first argument), population aggregate true, population red arm false. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HL72ndT9xg2dZEb6J8a2Gr
…t and counts as Nat
Three findings from review 58087, all upheld.
THE FAIL-CLOSED ONE. highest_quality seeded its fold with a helper that
MANUFACTURED a verdict when the list was empty: ServingCandidateRejected with
model_name "none" and DoesNotFitMemory { 0, 0 }. The review's observation that the
only caller can never reach it is precisely what made it dangerous rather than
harmless -- it is a synthetic row a later consumer would read as a real model
rejected for not fitting a real zero-byte capacity, and nothing in the type marks
it invented. A failure arm must refuse, never fabricate.
std carries no NonEmptyList, and hand-rolling one here would be a workaround for a
missing substrate type rather than a fix. So the fold is made TOTAL by answering
with an option instead of a value: highest_quality returns CandidateVerdict?
seeded with `none`, and the caller turns Absent into the honest
NoCandidateAdmissible. The fabricated constructor is deleted outright, and the
now-redundant `length(admissible) == 0` test goes with it -- emptiness had two
representations and now has one.
PARALLEL REPRESENTATION. DeclaredContext.max_positions and
YarnScaling.original_positions were bare Int while choice.dag typed the same
concept as TokenCount: one fact, two spellings, free to drift. Both are TokenCount
now. factor stays Int, correctly, as a dimensionless ratio.
COUNTS. ParameterCounts (all three), NarrowedFrom.rejected_count,
CompletenessAnswerable.member_count and ShardedWeights.shard_count admitted
negatives; all are Nat. quality_rank stays Int deliberately -- it is an ordinal,
not a magnitude.
Evidence after the change: selector aggregate true, selector refusal arm true
(this is the path the fabricated verdict used to sit on, so it is now executed
rather than merely unreachable), population aggregate true, population red arm
false.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HL72ndT9xg2dZEb6J8a2Gr
CI cause, not a new feature. Adding TokensPerSecond to std/measure.dag -- which review 58079 correctly required, so that prefill throughput consumes a measure carrier instead of a bare Int -- has a mandatory second half: the seed compiler builds from src/v1/stage0/src/std_measure.rs, a GENERATED mirror of that module. The .dag carried the type and the mirror did not, so the bootstrap build broke. That is why required-witnesses-build failed alongside the floor rather than the floor alone. The hunk is the regen actuator's own output installed verbatim, not hand-written into a file whose header says do-not-edit: claim_executor --required-regen --source-root dag --source-root src/v2 reported `FAIL generated surface drift: std_measure.rs` and named exactly one divergent file. Its candidate tree differs from the checked-in mirror by one hunk of thirteen lines -- the type alias and the two accessors -- at the offset the generator chose. WHY LOCAL VERIFICATION COULD NOT HAVE CAUGHT THIS, recorded because the blind spot is reusable rather than incidental. Every witness in this branch runs against a freshly compiled compiler, never against the mirror, so green-by-execution on a fresh build is structurally silent about seed drift. A change that touches a std module needs the regen check inside the authoring loop; discovering it from a required lane twenty minutes later is the expensive path. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HL72ndT9xg2dZEb6J8a2Gr
…rom what it happens to know
Two independent reviews approved the previous shape. Both missed that
choose_serving_candidate would answer confidently in two cases where no answer
exists, which is the failure this whole module was built to remove.
UNKNOWN != REJECTED. A candidate whose KV cost was never measured has not FAILED
the memory test; it has not taken it. The old shape had two outcomes -- chose, or
nothing admissible -- so an unmeasured candidate had nowhere to land, and a fully
measured but worse candidate would win by default with an answer indistinguishable
from one where it genuinely won. CandidateVerdict gains
ServingCandidateUnanswerable { missing: MissingFact }. For that arm to be
REACHABLE rather than decorative the measurements had to become genuinely absent,
so kv_per_token and measured_prefill_rate are optional; a zero sentinel would have
been the same fabrication deleted in the previous commit.
THE ANSWERABILITY RULE. ChoseCandidate is legal only when nothing unresolved could
change the optimum, because choosing the best ANSWERABLE candidate is not choosing
the best candidate. Any unresolved candidate now yields SelectionUnanswerable
{ UnresolvedCandidateCouldWin }, and the remedy it names is a measurement rather
than a purchase.
NO MANUFACTURED CROSS-RELEASE ORDER. Preference is ordinal WITHIN one release, so
two admissible releases have no join and a maximum over them does not exist.
Returning one anyway invents an ordering the inputs never contained. That case now
refuses with CrossReleaseQualityOrderAbsent. The practical consequence is that this
selector will NOT choose between DeepSeek V4 Flash and Qwen3.6 today, which is
correct: the missing input is operator judgment on real work, not another
measurement or more hardware.
Nat SUBTRACTION IS TOTAL (review 58094, non-blocking). allocatable and
firmware_carveout answered by subtracting unordered operands. Clamping to zero
would be the absorbing-fallback shape -- an impossible machine would quietly become
a machine with no room and every candidate would be rejected on memory for a reason
that was never true -- so both answer with an option and a malformed capacity
surfaces as unanswerable, a fact about the INPUTS rather than a verdict about any
candidate.
distinct_release_count uses the corpus fold/any/append idiom rather than importing
guarantee_measurement's distinct_string_count: std carries no distinct, and reaching
into a measurement authority for a list primitive from a model-selection module
would invert the layering.
Evidence: w_all_serving_choice_claims_hold true, w_all_answerability_claims_hold
true, w_without_the_context_floor_the_higher_quality_build_wins true (the control
separating a quality-maximizing selector from one returning its first argument),
w_at_the_400k_floor_the_two_bit_build_is_chosen true. The four new witnesses each
CONSTRUCT the state and observe it -- unmeasured, unresolved-blocks-choice,
two-releases-refuse, malformed-capacity -- so no Unanswerable arm is a decoration.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HL72ndT9xg2dZEb6J8a2Gr
briansrls
left a comment
There was a problem hiding this comment.
REQUEST_CHANGES at exact head 58e5423.
Accepted: the conservative answerability rule (any unresolved => SelectionUnanswerable) is safe even though it refuses more often than the stronger dominance-aware form. The new unanswerable arms are reachable and materially repair the prior best-answerable defect. The malformed-capacity option/refusal is also the right shape.
Blocking findings:
-
declared_contextis still being consumed as if it qualified the context floor.publication.dagcorrectly says DeclaredContext is a CLAIM and never a capability, and that context qualification comes from observed retrieval. Butevaluate_candidateadmits whenevercandidate.declared_context >= constraints.context_floor. That reintroduces the exact specification-without-execution rung error this PR is otherwise removing. Either selection must consume a semantic-context qualification/receipt (or an already-qualified population), or it must not decide the context axis yet. -
Latency remains a scalar candidate property (
measured_prefill_rate) and a scalar floor. The measured DeepSeek curve is depth-dependent (253@160k, 197@255k, 135@400k), and the operating regime is also causal: fresh-session prefill and warm continuation are different facts. Deferring orthogonal latency qualification to the population layer is acceptable only if this selector stops using a weaker scalar in its place. Current code still does, so the deferral is not neutral. -
NodeCapacityremains a second live-capacity authority.allocatable = kernel_visible - runtime_overheadis used to decide fit, while fabric already owns allocatable supply through SupplierOffer plus active grant reservations/conservation. A static hardware observation carrier is fine; a second selector-local answer to how many bytes are currently spendable is not. Feed selection from the fabric-qualified/remaining supply relation instead of recomputing live allocatable capacity here. -
Cross-release identity is implemented as
model_name: NonEmptyStr. This PR already definesReleaseIdentityprecisely because family/name/revision/publisher are distinct.distinct_release_countover model-name strings can merge distinct revisions/publishers or split aliases of one release. QuantizedCandidate must carry the typed release identity (or another exact release key), and the cross-release refusal must use that authority. Relatedly,quality_rankis still a bare Int rather than a within-release-scoped carrier, so the type itself does not prevent accidental cross-release comparison outside this function. -
Equal within-release quality ranks currently manufacture an ordering from input order.
highest_qualitykeeps the earlier candidate on ties while its comment says the result is stable under re-ordering; those statements are opposite. Either equal-rank survivors need an explicit deterministic policy/tie-break independent of list order, or selection must refuse as unanswerable/ambiguous.
Latency ruling: YES, fresh-session prefill and steady-state warm continuation need separate qualification regimes; restored/re-entered sessions should be a third regime if parked-state restore is part of serving. A realization may legitimately be fresh-session-latency-refused and warm-continuation-qualified. Warm qualification must bind to positive prefix/session-reuse evidence on the production path, not merely a fast second request.
Please keep #9897 unmerged until these are repaired or the selector is deliberately narrowed so it no longer claims to decide the deferred axes.
…forbids it The floor refused at 58e5423 on one row: required-floor: FAIL ...w_red_arm_forbidden_edge_must_evaluate_false returned Bool(false) verdict=FloorRefused unexpected_failures=1 The tree compiled; the row was authored to return false on purpose. The floor's only arm for that is v2.workflow.floor_expected_red, and that roster's own header rules it out -- "a row belongs here only while someone is fixing it", "an identity sitting here indefinitely is a defect nobody owns wearing a receipt". A permanent discriminating control is not debt, so enrolling it would have made the roster the skip list it is carefully not. Restated in the corpus idiom instead -- a positive assertion over a red input, the shape of ..._wrong_fixture_refuses_holds. w_population_membership_discriminates_in_ both_directions asserts a release the root population does not carry is not held, and the one it does carry is. That keeps what the old row was after (population_holds is not constantly true, so the claims above it are not vacuous) while reaching a terminal verdict. The forbidden edge itself was never carried by that row: RED 1 asserts it directly. Also removes three byte-identical declarations appended at the tail of choice.dag during the tie work -- verdict_is_unanswerable, distinct_release_count, verdict_model_name -- which the resolver refused as "a second declaration of one name silently replaced the first", and completes the MissingFact rename in the two serving-choice match sites that still listed the pre-split arm. Verified by execution, all seven returning true against the deduped tree: w_all_population_claims_hold, w_population_membership_discriminates_in_both_directions, w_all_serving_choice_claims_hold, w_all_answerability_claims_hold, w_tied_quality_ranks_refuse_in_both_roster_orders, w_a_declared_context_never_qualifies_without_a_retrieval_receipt, w_the_higher_quality_candidate_wins_when_both_are_admissible Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HL72ndT9xg2dZEb6J8a2Gr
…scope structural Side-chat review 5075343704 blocker 4: the cross-release wall was keyed by a model_name string and quality_rank was a bare Int whose scope lived in whichever function remembered not to compare it across releases. A name key is wrong in both directions -- two revisions of one family share a name and merge, one release named two ways splits -- and this PR had already minted an exact ReleaseIdentity for precisely that reason, so counting anything else was a second, weaker key for a fact the identity already decides. QuantizedCandidate and both refusal verdicts now carry it. The ordinal fix is the load-bearing half. WithinReleaseQuality pairs a rank with the release it is ordinal under, compare_within_release returns Absent across releases, and quality_maximum folds through it -- so the cross-release refusal now comes from the comparison having no arm that produces a number, not from the caller checking a distinct count first. That pre-check was validation standing where construction was available (DESIGN 5): satisfiable by editing the caller while the comparison still lied. distinct_release_count survives only to fill the diagnostic. Also drops the `0 - 1` rank verdict_quality returned for non-admissible candidates. That is an invented ordinal: it places a rejected candidate below every real rank on a scale it was never measured on, and collides outright with a release that ranks its own artifacts from -1. Absent instead. release_identity_equal is hoisted to gunbc.model.publication, which owns the identity. population.identity_in had inlined the three-field comparison, so a fourth component added to ReleaseIdentity would have left that consumer silently comparing three. Two things the compiler forced, both of which improve the result: `release` is a RESERVED KEYWORD in every identifier position, not only in a module path, so `release: ReleaseIdentity` refused the whole module index. The field is `identity:`, which is what ModelRelease in publication.dag already calls it. An annotation indented inside a coproduct body is a parse error; annotations are module-item grain. Review 58106 (approving, nit-tier): CapacityMalformed carried a `detail: String`. There is exactly one way a NodeCapacity is malformed, so the string could hold only one value and restated what the constructor says. Payload dropped; the sentence is rendered by missing_fact_wire. Verified by execution, all seven returning true, unfiltered with per-witness rc: w_all_population_claims_hold, w_population_membership_discriminates_in_both_directions, w_all_serving_choice_claims_hold, w_all_answerability_claims_hold, w_tied_quality_ranks_refuse_in_both_roster_orders, w_a_declared_context_never_qualifies_without_a_retrieval_receipt, w_the_higher_quality_candidate_wins_when_both_are_admissible The preceding filtered run of the same set was a FALSE PASS: a module-index refusal prints no line matching "returned|error", so an allowlist grep rendered a tree that did not parse as five clean headers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HL72ndT9xg2dZEb6J8a2Gr
|
Re: review 58100 — the predicate/walker finding. I looked for the rule this is measured against and could not find it. It is not in DESIGN.md, and it is not a row in I also measured the population rather than arguing about it. Re-derive with: The flagged shape is the corpus default, including inside The substantive halves of the review I have taken. — sent from eager-pike-541 |
briansrls
left a comment
There was a problem hiding this comment.
REQUEST_CHANGES at exact head b825cc3.
Accepted on this head: blocker 1's comparison fix (declared context no longer admits the floor), blocker 4 (ReleaseIdentity + structurally scoped WithinReleaseQuality), blocker 5 (ties refuse independent of roster order), and the conservative any-unresolved => SelectionUnanswerable rule.
I WITHDRAW prior blocker 3 AS STATED. I treated supply.dag's prose ('remaining capacity is derived as allocatable bound minus active grant claims') as if an executable byte-capacity authority existed. It does not: an exact search finds that statement only in prose, while ExecutionGrant carries opaque ReservationRefs and the fabric reconciliation explicitly calls resource-conservation evidence type-erased. There is therefore no fabric remaining-capacity symbol for this PR to consume. product.node_power_envelope.node_memory_capacity is the nominal DIMM/population sum; it may source NodeCapacity.nominal where the physical inputs match, but it must NOT source kernel_visible. kernel_visible remains an independently observed OS-visible fact.
Remaining blockers:
-
LATENCY REGIME + REALIZATION PROVENANCE. The depth-indexed PrefillObservation repair is good, but PrefillObservation still has only {depth, rate}, ServingConstraints has no regime, and the observation sits on QuantizedCandidate. A FreshSessionPrefill measurement can therefore be mixed with a future WarmContinuation measurement and satisfy the same floor. More generally, semantic_context_verified_to, prefill_observations, kv_per_token, and runtime_overhead are realization facts: the same release/quant can behave differently under another runtime digest/config/executor/KV representation. A regime discriminator on observation plus a required regime on the constraint is necessary, but sufficient for this PR only if the selector also binds those observations to the exact serving realization it is selecting and mismatched/absent regime evidence => Unanswerable. The selector does NOT need to validate cache reuse itself. That evidence can live in the later qualification layer, but WarmContinuation/RestoredSession must not become authorable from naked numbers before that evidence carrier exists. Safe narrow choices: (a) consume a realization-bound/sealed qualification ref and let the later layer mint warm/restore receipts, or (b) keep only fresh-prefill constructible here and let warm/restore constraints refuse until the evidence layer lands, or (c) remove latency admission from this selector for now.
-
THE DEFERRED POPULATION LAYER CANNOT CURRENTLY CARRY REALIZATION QUALIFICATION. ModelPopulation.members is List, yet RuntimeCompatibleStage, SingleNodeFeasibleStage, MultiNodeFeasibleStage, ContextQualifiedStage, and CodingQualifiedStage are explicitly described as verdicts about a REALIZATION. DeepSeek IQ2 vs IQ3 is the concrete falsifier: one realization can fit/qualify at 400k while another realization of the same release does not. Narrowing by ReleaseIdentity erases that distinction. Before using population as the home for semantic/latency evidence, either introduce a realization-grain population/key for those stages or delete/defer those stages from the release population until that carrier exists.
-
STALE HARDWARE-REMEDY CLAIM. choice.dag still says 'If every candidate is bound by memory, more nodes help' and describes rejection axes as partitioning into 'buy hardware'. That was explicitly withdrawn: a memory rejection proves only that more usable memory might help; whether another node helps depends on a distributed realization/topology/runtime. The accepted mechanism is the counterfactual query over current supply vs current+proposed supply. Remove the false implication before merge.
Capacity note after withdrawing blocker 3: the selector-local kernel_visible - runtime_overhead value is presently a STATIC realization/planning ceiling, not a live remaining-capacity authority. Please name/document it that way; do not claim it is current spendable memory. If runtime_overhead can differ by selected runtime/config, bind/move it with the serving realization rather than the node.
Latency ruling remains: FreshSessionPrefill, WarmContinuation, and RestoredSession are distinct regimes. Reuse evidence and concurrency conditions belong to the observation/qualification producer, not to the selector. The selector's job is exact regime/realization matching and fail-closed absence.
Side-chat blocker 2, to the bar they set: the selector must make regime mismatch
unanswerable rather than let a fresh-prefill receipt satisfy a warm constraint.
ServingRegime = FreshSessionPrefill | WarmContinuation | RestoredSession is now a
field on PrefillObservation and on ServingConstraints, and prefill_rate_at_floor
filters on regime equality as well as depth. There is no widening arm: a candidate
holding only fresh receipts, asked about a warm floor, returns Absent, which
evaluate_candidate turns into PrefillRateUnmeasuredAtFloor { floor, regime } and so
into Unanswerable. It has not failed the test, it has not taken it.
The substitution is flattering in exactly one direction, which is why the arm cannot
exist rather than merely being unused: warm numbers are the fast ones, so letting a
warm receipt answer a fresh question admits a realization that cannot serve a cold
request at the floor. This fleet measures 2963 s for a cold 400k prefill against
seconds for an appended turn -- three orders of magnitude, not measurement noise.
Every observation in the fixture is labelled FreshSessionPrefill, which is what they
are: every latency number measured on the nodes was a cold prefill. RestoredSession
is enumerated with no evidence behind it deliberately, so a future observation of a
session reloaded after eviction or slot reassignment must say which of the three it
is rather than defaulting into WarmContinuation.
Two witnesses, on the helper and on the selector, so the refusal is not merely a
property of a filter that never fires. The first carries its own positive control:
the SAME observation list does answer at FreshSessionPrefill and yields 135 tok/s, so
the Absent is not satisfied by a filter that rejects everything.
NOT addressed here, and still owed: the reuse-evidence receipt (same realization and
cache identity, known reused-token population) and the request-concurrency condition.
Those are observation-side obligations a selector cannot discharge; this change only
ensures a warm claim cannot be made without them.
Review 58113 (approving, nit-tier): distinct_release_count carried two doc blocks --
the ReleaseIdentity rewrite left the pre-rekey annotation stranded above its
replacement, so one explained a name key and one an identity key. Stale one deleted.
Verified by execution, all seven returning true, unfiltered with per-witness rc:
w_all_serving_choice_claims_hold, w_all_answerability_claims_hold,
w_a_fresh_prefill_receipt_does_not_answer_a_warm_continuation_floor,
w_a_warm_floor_over_fresh_only_receipts_is_unanswerable,
w_tied_quality_ranks_refuse_in_both_roster_orders, w_all_population_claims_hold,
w_population_membership_discriminates_in_both_directions
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HL72ndT9xg2dZEb6J8a2Gr
briansrls
left a comment
There was a problem hiding this comment.
REQUEST_CHANGES at exact head 911b4a3.
Accepted on this head: the regime mismatch itself is modeled correctly. ServingRegime is present on PrefillObservation and ServingConstraints; prefill_rate_at_floor filters by exact regime and depth; the helper witness has a same-roster Fresh positive control; and the selector propagates a Warm-over-Fresh mismatch to PrefillRateUnmeasuredAtFloor => Unanswerable. That closes the prior cross-regime borrowing defect.
Blocker 3 is WITHDRAWN definitively. I re-ran the repository-wide search: product.fabric has prose saying remaining capacity should be derived, but no executable byte-valued remaining-capacity function/ledger to consume. ExecutionGrant's resource reservations are opaque ReservationRefs. product.node_power_envelope.node_memory_capacity is a nominal DIMM/population sum and may source NodeCapacity.nominal, not kernel_visible. I will not require this PR to consume a nonexistent fabric symbol.
Remaining blockers:
-
REGIME LABEL != EVIDENCE WALL. PrefillObservation remains a freely authorable record { regime, depth, rate }, and ServingRegime exposes WarmContinuation and RestoredSession as ordinary constructors. Therefore a caller can still write PrefillObservation { regime: WarmContinuation, ... } from naked numbers with no cache/session-reuse evidence. The new filter prevents BORROWING a Fresh measurement into Warm, which is good, but it does not make an unsupported Warm claim unwritable. Since the reuse-evidence producer is intentionally deferred, narrow this PR structurally: either make PrefillObservation sealed and expose only a FreshSessionPrefill constructor now, leaving Warm/Restored constructors for the evidence layer; consume an opaque/sealed realization-bound qualification minted upstream; or remove Warm/Restored observation construction until that layer lands. The selector itself does not need to validate reuse.
-
REALIZATION GRAIN IS STILL MISSING. QuantizedCandidate carries release identity + quant label, while semantic_context_verified_to, prefill_observations, and kv_per_token are runtime/config/executor-dependent realization facts. More importantly, model.population says RuntimeCompatibleStage, SingleNodeFeasibleStage, MultiNodeFeasibleStage, ContextQualifiedStage, and CodingQualifiedStage are verdicts about a REALIZATION, but ModelPopulation.members is List and narrow_population admits only ReleaseIdentity. DeepSeek V4 IQ2 vs IQ3 is the concrete falsifier: one realization can fit/qualify while another of the same release does not. A release-only carrier cannot preserve that distinction. Either introduce a RealizationIdentity/RealizationPopulation for those stages, or narrow ModelPopulation to release-grain stages and defer realization stages until the proper carrier exists.
-
STALE DERIVATION CLAIM. choice.dag still states that if every candidate is memory-bound, more nodes help, and describes rejection axes as partitioning into 'buy hardware' remedies. We already withdrew that inference: a memory rejection proves only that more usable memory could remove the rejection; another node helps only if the realization/runtime/topology can use it. Replace that prose with the counterfactual rule (rerun the same selector over current supply vs current+proposed supply) or delete it.
I do NOT require the reuse receipt or concurrency-condition producer itself in this PR if unsupported Warm/Restored claims are structurally unmintable here. Once those three items are repaired, the original blocker set is otherwise closed.
… buy-nodes rule is gone Three narrow fixes, each closing a gap the reviewer located on 911b4a3. REGIME PROVENANCE. Discrimination was solved; provenance was not. `PrefillObservation` carried a three-arm regime field, so `{ regime: WarmContinuation, rate: 500 }` was an ordinary writable record -- the fast number, mintable by anyone, with no reuse receipt behind it. The carrier is now `FreshPrefillObservation` with no regime field at all: the only regime whose provenance needs no receipt is the one that measures itself. Warm and restored floors return `none` structurally rather than by filtering, so the refusal survives any edit short of building the receipt, whose carrier is named as the trigger. POPULATION GRAIN. `PopulationStage` enumerated five realization verdicts over members typed `List<ModelRelease>`. This repository's own measurement falsifies that keying: DeepSeek V4 Flash 0731 at IQ2_XXS is single-node feasible at a 400k floor and the same release at IQ3_S is not, and a release-keyed stage answers "DeepSeek V4 is in" for both. The stages are removed rather than renamed, with RealizationIdentity/RealizationPopulation named as the carrier they belong to. RED 3 now witnesses the stronger invariant it was always reaching for: no narrowing, empty or otherwise, reaches back into its parent. STALE DERIVATION. Two sites still said a memory-bound census means more nodes help. That rule was withdrawn: a memory rejection proves only that more USABLE memory would remove it, and whether another node supplies memory this realization can use is a distributed- runtime and topology fact the axis does not carry. The hardware question is the counterfactual -- choose(current) against choose(current + proposed) -- and a rejection census is its input, not its answer. Every witness re-run and passing at identity grain: population, serving choice, answerability, and the tie refusal.
briansrls
left a comment
There was a problem hiding this comment.
APPROVE at exact head 14fb18a.
The three remaining blockers from review 5076550603 are closed:
-
Warm/restored provenance: FreshPrefillObservation has no regime field, so unsupported WarmContinuation/RestoredSession observations are not merely filtered out; they are unrepresentable. prefill_rate_at_floor answers those requested regimes with Absent structurally, and the selector propagates that to PrefillRateUnmeasuredAtFloor/Unanswerable. PrefixReuseReceipt is correctly named as the next-rung trigger rather than fabricated in this selector.
-
Population grain: ModelPopulation is now honestly release-grained. PopulationStage contains only OpenWeightStage and LocallyPackagedStage; realization-level runtime/fit/context/coding stages were deleted, with RealizationIdentity/RealizationPopulation named as the carrier required before those verdicts return. The DeepSeek IQ2/IQ3 falsifier is preserved in the rationale.
-
Hardware-help derivation: the stale
memory-bound => more nodes helpclaim is gone. Rejection axes now state which input would have to move, while hardware benefit is explicitly a counterfactual selection over current supply versus current+proposed supply.
Prior blocker 3 concerning a duplicate fabric live-byte remainder remains withdrawn: there is no executable byte-valued fabric remainder authority for this PR to consume.
I do not require the allocatable rename in this PR. That name should move together with the later decision about runtime_overhead/realization authority rather than as an isolated terminology patch.
Exact-head witness CI was still in progress when I submitted this review, so this approval is technical authorization of the tree, not a claim that the required CI has completed. Merge should remain gated on the repository's ordinary exact-head CI/attestation requirements.
briansrls
left a comment
There was a problem hiding this comment.
EMERGENCY STOP / REQUEST_CHANGES at exact head 14fb18a. My prior approval on this head is withdrawn and superseded. Do not merge.
The author has produced direct execution evidence that falsifies the fixture's central claimed measurement: kv_per_token = 88,064 bytes is labeled "exact rather than estimated" and drives the showcase 400k-context memory rejection, but real loads on spark-a3ee under Ollama 0.32.9 contradict it. IQ3_S loaded at num_ctx=400,000, reported about 116.97 GB VRAM, and generated a token, so it fits beneath the fixture's 127.66 GB ceiling. IQ2_XXS also loaded at num_ctx=1,048,576 with about 88.9 GB reported; the claimed 88,064 B/token would imply roughly 92.3 GB of KV alone plus 86.7 GB of weights, which is incompatible with that successful load on a 121.69 GiB node.
Therefore the headline witness is currently wrong: at the 400k floor IQ3_S is not memory-rejected by this evidence. Both same-release realizations are admissible, so the established within-release quality ordering should select IQ3_S rather than IQ2_XXS.
This does not reopen the selector-logic repairs already accepted: evaluation order, fail-closed answerability, regime split, release-population grain, and counterfactual hardware rule remain intact. The blocker is the fixture's publication-layer rung error: an architecture-derived declaration was allowed to masquerade as an observed qualification.
Before re-review, kv_per_token must be sourced as an instrument- and realization-bound observation (including the named node/runtime/method), any architecture-derived figure must be removed or made structurally unable to qualify the floor, and the 400k witnesses must be re-derived from the observation rather than re-asserted. No substantive re-review is being attempted in this stop review.
briansrls
left a comment
There was a problem hiding this comment.
AMENDMENT TO EMERGENCY STOP / REQUEST_CHANGES at exact head 14fb18a. This supersedes the KV-repair prescription and the claimed final winner in review 5080033756. The stop remains in force: do not merge. This is not a re-review of any repair.
- THE LINEAR FORM, NOT JUST 88,064, IS FALSIFIED.
Direct Ollama 0.32.9 /api/ps size_vram observations for the exact IQ2_XXS realization on spark-a3ee are:
- 131,072 -> 86,532,465,622 B
- 262,144 -> 87,064,355,798 B
- 400,000 -> 87,692,389,907 B
- 1,048,576 -> 88,865,253,620 B
The successive segment slopes are exactly 4,058 B/token, approximately 4,555.73 B/token, and approximately 1,808.37 B/token. Therefore no corrected kv_per_token scalar can repair weights + kv_per_token * context * hot_sessions; replacing 88,064 with a better-looking constant would preserve the same unsupported extrapolation.
Required shape before re-review:
- Remove
kv_per_token,kv_footprint, and the linearresident_footprintcalculation from the qualifying fit path. - Consume an observed resident-footprint fact at or beyond the requested context, bound to the exact realization and the observation conditions: artifact/model identity, runtime release/config, node, instrument/method, and concurrency/session condition.
- A shallower observation, a realization/runtime mismatch, or absent evidence must yield a typed unanswerable result. Do not interpolate or extrapolate between depths.
- Do not recover concurrency by multiplying a one-session reading. A one-session observation cannot answer a
hot_sessions > 1constraint; concurrency must be observed/bound or the question must refuse. - Return/carry the qualifying observation (including its actual depth), not only a detached byte scalar, so the admission remains auditable.
- Architecture dimensions may remain as a declaration only if there is structurally no path from them to memory qualification.
The exact 400k IQ3_S observation (about 116.97 GB) establishes that its MEMORY axis fits under the fixture's 127,660,151,296-byte ceiling. It invalidates the existing memory rejection and IQ2-by-memory witness.
- THE MEMORY RESULT REVERSES; THE FINAL SELECTOR ANSWER DOES NOT YET BECOME IQ3_S.
My prior stop review overstated the consequence by saying both realizations were already admissible and IQ3_S should win. On this exact head, build_iq3_s still has semantic_context_verified_to: none and fresh_prefill_observations: []. Loading a runner at num_ctx=400000 and generating one token establishes resident fit; it does not establish semantic retrieval at 400k or fresh-prefill performance at that floor.
After the footprint repair, the honest result over the currently evidenced installed roster is therefore SelectionUnanswerable { UnresolvedCandidateCouldWin }: the higher-quality IQ3_S no longer loses on memory, but remains unqualified on later axes and could win when measured. The real fixture may claim IQ3_S as the selected winner only after those independent receipts exist. A separate declared/synthetic fully qualified same-release pair may continue to witness that the higher ordinal wins when both candidates are admissible.
The previously accepted selector logic remains accepted: unknown is not rejected, no best-effort fallback, fixed evaluation order, exact regime refusal, release-grain population, within-release quality, and counterfactual hardware reasoning.
…rchitecture constant
The fixture declared kv_per_token = 88,064 bytes and called it "exact rather than
estimated" because it was derived from the architecture -- 43 blocks, one latent KV head,
key/value 512. It was a derivation wearing the label of a measurement, in the fixture for
the module whose subject is refusing exactly that substitution.
Loaded on spark-a3ee and read back through /api/ps size_vram, one build residents at
86,532,465,622 / 87,064,355,798 / 87,692,389,907 / 88,865,253,620 bytes for 131,072 /
262,144 / 400,000 / 1,048,576 tokens. Segment rates 4,058, 4,556, 1,808 bytes per token.
THE FUNCTION IS FALSIFIED, NOT ITS COEFFICIENT, so no replacement scalar is authored: an
averaged slope or a fitted curve would only manufacture a second unobserved answer with
better provenance. The constant also fails alone -- at 88,064 B/token a 1,048,576 window
is 92.3 GB of KV, exceeding the whole node beside 86.5 GB of weights, and that
configuration loads and serves.
So kv_per_token, kv_footprint, resident_footprint and the weights field are deleted, and
fit is established from a ResidentFootprintObservation carrying realization, runtime,
node, instrument, depth, concurrency and resident bytes. Deeper qualifies shallower and
never the reverse, the tightest qualifying bound is returned, and CONCURRENCY IS NOT
RECOVERED BY MULTIPLICATION -- a one-session reading cannot answer a two-session demand,
so that is Unanswerable rather than scaled. The whole observation is returned rather than
its byte count, so an established fit stays locatable.
TWO RESULTS REVERSE. The higher-quality IQ3_S is not rejected on memory at a 400k floor;
it residents in 116,970,000,000 B against 127,660,151,296 B allocatable. And fitting is
not winning: it has no retrieval receipt and no prefill observation, so it is unresolved,
and the real roster now REFUSES rather than handing the 2-bit build a default win.
The evaluation-order annotation is corrected too. It claimed memory fit was decidable from
declared facts, which is why ordering it first cost nothing; fit is now evidence-bearing
and the memory arm can itself refuse, so the order buys refusal economy and not
decidability.
Also recorded: a fold accumulator bound by Present { value: x } does not carry its type to
a field access, so the comparison is a named binary operation. That is a language-layer
gap, noted where it bites rather than routed around silently.
The 1,048,576 desired-context change is a FLEET change: it alters what the two Sparks are configured to serve. This PR's subject is the model selection authority. They were riding together only because I made both edits in the same hour, which is not a reason. It lands on its own branch with the rung drop it needs -- the measurement behind 1M is bound to one node, one IQ2_XXS artifact, one runtime and one concurrency condition, and it is being applied to two hosts and every realization on the roster. That declaration does not belong in a diff about ReleaseIdentity.
briansrls
left a comment
There was a problem hiding this comment.
HOLD / REQUEST_CHANGES at exact head 8917397. Do not merge. This is an exact-head stop, not a substantive re-review of the footprint repair.
The earlier a70a8ac repair head has been superseded by this head. The author has identified a remaining accounting-domain defect in the current repair: an Ollama /api/ps size_vram observation is named as a generic ResidentFootprint and compared directly with allocatable(capacity), which is derived from /proc/meminfo MemTotal minus declared overhead. Ollama 0.32.9 reports size_vram as runner GPU-buffer usage parsed from llama-server output; that does not, by itself, establish an accounting bridge to OS-visible memory or authorize exact headroom subtraction. The type and field names currently assert that bridge rather than prove it.
Before re-review:
- preserve
/api/pssizeandsize_vramas the exact instrument-reported quantities rather than renaming either to generic resident memory; - carry successful load/generation as its own positive fit claim;
- retain exact realization, runtime/configuration, node, context depth, and concurrency provenance;
- do not emit exact headroom or compare across the Ollama-buffer and OS-memory domains until an explicit authority establishes that accounting relation for the observed configuration;
- where
size == size_vramor full offload matters, carry it as an observation/receipt rather than infer it from the platform.
The prior selector-logic acceptances remain intact, including the corrected real-roster result: after memory fit reverses, IQ3_S is still unqualified on semantic-context and fresh-prefill axes, so the honest answer is SelectionUnanswerable { UnresolvedCandidateCouldWin }, not an IQ3_S winner.
The reported removal of serving_desired.dag and the fleet-global context/concurrency work are outside this stop review and are not being adjudicated here. Re-review remains deferred until the author explicitly requests it.
…ps subtracting one from the other The repair I pushed for the falsified KV constant introduced a subtler version of the same error, and review caught it before it decided anything. The carrier was named ResidentFootprintObservation with a field `resident`, and that number was compared against allocatable(capacity) -- kernel-visible memory from /proc/meminfo less a declared overhead. /api/ps reports `size` and `size_vram`, which Ollama documents as total and GPU BUFFER usage parsed from its runner's logs. That is not an operating-system resident set. The comparison spanned two accounting domains, so the headroom it produced was manufactured, and the name asserted the bridge that no measurement had crossed -- the same move as labelling a derived constant "exact", one layer up. So fit is no longer computed and no longer compared. IT IS EXECUTED. The carrier is OllamaRunnerMemoryObservation, capturing both reported buffer figures without renaming either into an OS quantity, plus the realization, runtime, runtime configuration, node, instrument, depth, concurrency, and whether generation SUCCEEDED. A successful load is a positive fit under exactly the conditions that held; a failed one is an observed non-fit; no observation is Unanswerable. Node equality joins the qualification, because a fit established on one machine says nothing about another. DELETED WITH THE COMPARISON: NodeCapacity, allocatable, firmware_carveout, CapacityMalformed, and the two witnesses whose subject they were. The three-grain distinction between vendor nominal, kernel-visible and live-available is real and has already cost one wrong decision, but it belongs to an authority that owns node facts, not to a selector that can no longer compare against them. Keeping the carrier so its witness had something to test would be a check keeping its own subject alive. The starved-node rejection witness moved for the second time. It first used an absurd context floor, which became Unanswerable once fit needed evidence; then a starved NodeCapacity, which is gone with the subtraction. What remains is the only non-fit the instrument can report: a configuration attempted on the node that did not produce a token. Every witness aggregate re-run and passing: serving choice, answerability, tie refusal, population.
briansrls
left a comment
There was a problem hiding this comment.
REQUEST_CHANGES at exact head c5cb567. This is the requested substantive re-review. Do not merge.
Accepted from the accounting repair:
kv_per_token, the linear KV/weights product, the OS-memory subtraction, exact headroom,NodeCapacity,allocatable,firmware_carveout, andCapacityMalformedare gone from the selector. I agree with deleting them rather than preserving a decorative carrier in this PR./api/pssizeandsize_vramare now named as Ollama-reported buffer quantities, not OS resident memory.- depth, node, and concurrency participate in the current helper; shallower evidence does not answer a deeper floor; no multiplication reconstructs concurrency; the whole selected observation travels in the admissible verdict.
- the real roster no longer claims IQ3_S wins merely because its memory rejection disappeared. Its intended answer remains
SelectionUnanswerable { UnresolvedCandidateCouldWin }. serving_desired.dagis out of this PR.
The exact head still has the following blockers.
- CARRIED PROVENANCE IS NOT JOINED PROVENANCE.
runtime_memory_observation_at_floor filters only context depth, concurrent_sessions, and node. It never compares observation.realization, runtime, runtime_config, or instrument against the candidate or against a requested serving realization. QuantizedCandidate itself has only release identity plus a quant label; it has no exact runtime/configuration identity to compare with. Therefore an observation for another artifact, another runtime, another config, or another instrument can be placed in the candidate's list and qualify unchanged.
The real fixture demonstrates the weakness: the IQ2 realization is a mutable ...:latest string, IQ3 is another tag string, runtime is a prose version string, and runtime_config is prose. None is an immutable artifact/runtime/config identity, and none is read by the selector. Changing all four strings—or moving the IQ2 observation into the IQ3 list—leaves the current headline witnesses green.
The comment also says buffer comparisons are between readings from ONE instrument, but the code folds every qualifying observation regardless of instrument/runtime/config and compares their reported_total_buffer values. That same-domain property is asserted only in prose.
The realization-layer trigger has fired for this selector. It need not build the future whole population, but it does need an exact typed serving-realization/configuration key carried by the candidate and every qualifying observation, with equality enforced. The semantic-context and prefill evidence must join that same key as well; otherwise the selector can compose memory fit from runtime A, retrieval from B, and prefill from C and call the nonexistent combination admissible. Add negative controls for wrong artifact/realization, runtime, runtime config, instrument, and node.
generation_succeeded: BoolDOES NOT ESTABLISHDoesNotFitMemory.
A missing token can be caused by an invalid request, tokenizer/template failure, unsupported operation, cancellation, timeout, runtime crash, transport failure, or many other causes. The current code maps every false directly to DoesNotFitMemory. That is a diagnosis the observation did not make.
The carrier also requires /api/ps buffer fields. A load that fails before a runner exists may have no /api/ps row from which those fields can be read; requiring byte values forces the failed-load case to fabricate them. Conversely, a runner may load and report buffers successfully, then fail generation for a non-memory reason—in which case its memory fit was not disproved.
Split the outcome structurally. A successful live-runner memory observation may carry /api/ps buffer readings and a positive serve receipt. A failed attempt needs a separate carrier with the exact attempted configuration and a typed observed failure cause. Only an OOM/insufficient-memory-specific observation may become DoesNotFitMemory; other failures need their own axis or an honest unanswerable/refusal. The present fixture_load(..., succeeded: false) is not a valid memory non-fit witness.
tighter_boundCAN CHOOSE THE VERDICT BY BYTE MAGNITUDE.
The qualifying list includes both successful and failed observations. The fold chooses the smallest reported_total_buffer before evaluate_candidate inspects the outcome. Thus a smaller failed reading can hide a successful fit and reject the candidate, while a smaller successful reading can hide a failed receipt and admit it. The same problem occurs across different runtime configs and instruments because finding 1 leaves them in one fold.
Once bytes are no longer compared with capacity, "smallest buffer" is not evidence of fit and is not the selector's objective. A failed attempt at a deeper context also does not prove failure at a shallower floor where another configuration succeeded. Match the exact requested realization/configuration first and reconcile typed outcomes explicitly. If multiple incomparable or contradictory receipts remain, refuse or return the conflict; do not let buffer magnitude settle truth. Add both roster-order controls and mixed success/failure controls.
- THE REAL-ROSTER HEADLINE WITNESS DOES NOT PROVE ITS HEADLINE.
w_the_real_roster_refuses_while_the_higher_rank_is_unresolved accepts any SelectionUnanswerable and checks only that evaluations has length two. It does not require UnresolvedCandidateCouldWin, does not require unresolved_count == 1, and does not establish that IQ2 is fully admissible while IQ3 is specifically semantic-context-unanswerable. A cross-release refusal, tie refusal, or two unresolved candidates would satisfy it.
Pattern-match the exact cause and both evaluations. Likewise, the moved all-rejected witness should inspect two DoesNotFitMemory arms backed by qualifying typed memory-failure receipts rather than only checking list length.
- THE OLD IQ3_S NON-FIT CLAIM IS STILL PRESENT.
gunbc.model.population still says the repository's concrete realization-grain falsifier is: IQ2_XXS is single-node feasible at 400k while IQ3_S is not. The direct load measurement falsified that sentence. The release-vs-realization grain ruling remains correct, but this example must now say what is actually established—for example, IQ2 is qualified on the modeled axes while IQ3 has observed memory fit and unresolved later axes—or use another real falsifier. A known-false rationale cannot remain because its structural conclusion survived.
- GLOBAL COMPLETENESS IS ANSWERED FROM ONE PUBLISHER'S CATALOG.
This is a separate exact-head finding from the full re-review. coverage_spans_open_weight_universe returns true for any PublisherReleaseCatalog { publisher: _ }, and open_weight_completeness then answers for the entire population. A population containing DeepSeek and Qwen releases with only a DeepSeek catalog source is therefore declared complete. That is the same wrong-population error this module says it prevents.
Scope the completeness question to one publisher, or require complete source coverage for every publisher in an explicitly declared target universe. Add a mixed-publisher negative control; the current one-publisher positive fixture cannot distinguish the bug.
- THE POPULATION SUBSET CLAIM IS NOT STRUCTURAL YET.
The module says a narrowed population is not authored, is constructed only from a parent, and retains that parent. But ModelPopulation and PopulationProvenance are ordinary constructible records/variants, so an arbitrary stage/member/provenance combination is writable. NarrowedFrom retains only parent_stage, not the parent population or its member identity, and narrow_population accepts an arbitrary output stage, including backward or same-stage transitions. The tests exercise the helper; they do not make bypassing it impossible.
Either make the accepted population carrier opaque/constructor-bound with fixed legal transitions and a retained parent identity, or state the rung honestly as mechanical prevention and add an enrolled gate over every authored population. The current source cannot support the claimed structural guarantee.
Previously accepted selector properties remain accepted: unknown is not rejected, no best-effort fallback, warm/restored observations remain unwritable, release identity and within-release quality remain scoped, ties refuse, and the buy-nodes inference remains replaced by counterfactual selection. Exact-head CI does not cure the semantic counterexamples above.
…n attempt produced Findings 1-7 of the side chat's REQUEST_CHANGES. (1) Memory, retrieval and prefill evidence now join a ServingRealizationIdentity -- release x quantization x artifact x runtime x runtime-config. The fields were carried and never matched, so evidence from another artifact, runtime or configuration qualified unchanged, and memory fit from A plus retrieval from B plus prefill from C could be admitted as one candidate that never existed. Negative controls flip each component of the key in turn, plus the node, and require the evidence to stop answering; the positive control keeps them from being satisfied by a filter that rejects everything. (2) An unsuccessful attempt is no longer a memory verdict. RunnerAttemptOutcome splits served / refused-for-memory / failed-for-other-cause, and the reported buffers live only on the arm where a runner existed to report them -- so a pre-runner failure carrying byte figures is unwritable rather than discouraged. Only the typed insufficient-memory refusal rejects; every other failure is Unanswerable and names its cause. (3) Magnitude no longer decides between contradictory readings. The fold picked the smallest reported buffer BEFORE reading the success flag, so a small failed receipt beat a large successful one. Qualification is now an identity join first, and two qualifying attempts that disagree refuse -- exercised in both roster orders, with the large-buffer success alone still admitting. (4) The real-roster claim asserts the exact cause (UnresolvedCandidateCouldWin, unresolved_count == 1) and the exact verdict arms, not just the evaluation count. (5) The IQ3_S non-fit rationale is withdrawn from population.dag -- loading the build falsified it. The finer split the measurements DO establish replaces it, cited by naming the witness rather than transcribing verdicts. (6) Completeness is publisher-scoped. Answering Answerable for any publisher catalogue said one publisher's index enumerates the open-weight universe: a real count over the wrong population, the original error re-committed by the function written to prevent it. The universe-wide question has no answerable arm. (7) The population subset claim is rung-honest. narrow_population constructs a subset and now retains the parent roster and refuses backward or same-stage transitions, but ModelPopulation is an ordinary record -- the class is mechanically preventable, not structurally impossible, and its trigger is constructor privacy in the substrate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
|
All seven findings of the REQUEST_CHANGES are addressed at (1) The identity join. Instrument identity is deliberately not a join component. It names the procedure that produced a reading, not the subject the reading is about — two instruments reading one realization at one configuration should agree, and where they do not, that is finding 3's contradiction, which now refuses without consulting the instrument. Joining it would split one subject into two populations and make the disagreement invisible instead. (2) An unsuccessful attempt is not a memory verdict. (3) Magnitude no longer arbitrates. Qualification is an identity join first; a served and a memory-refused attempt at the same joined configuration refuse rather than being settled by a number. Exercised in both roster orders with a 99 GB success paired against a refusal, and with the same success alone still admitting as its positive control. (4) The real-roster claim asserts the exact cause ( (5) The IQ3_S non-fit rationale is withdrawn from (6) Completeness is publisher-scoped. (7) The subset claim is rung-honest: mechanically preventable, not structurally impossible, with constructor privacy in the substrate as the next-rung trigger.
— sent from eager-pike-541 |
briansrls
left a comment
There was a problem hiding this comment.
REQUEST_CHANGES at exact head eacbb2f. This is the requested substantive re-review. Do not merge.
Accepted on this head:
- I agree that
instrumentmust NOT be a component of the serving-subject join. It names how a reading was produced, not what was measured. Joining it would split one subject into instrument-specific populations and hide a same-subject disagreement. Carrying it for audit while reconciling outcomes across instruments is the right direction. RunnerAttemptOutcomecloses the Bool/counterfeit-buffer defect. Buffer quantities exist only onRunnerServed; onlyRunnerRefusedForMemoryreachesDoesNotFitMemory; a non-memory failure is honestly unanswerable.- Byte magnitude no longer directly chooses success versus refusal.
- The real-roster witness now fixes the exact top-level cause and count, and its companion claims establish the IQ2/IQ3 arm split.
- The known-false IQ3 memory rationale is withdrawn from
population.dag. - Publisher-scoped completeness and the permanently refusing universe arm close finding 6.
- I do not require a warm/restored reuse receipt in this PR while observations remain structurally fresh-only and non-fresh questions return Absent.
The following blockers remain.
- THE JOIN IS LIVE, BUT THE SUBJECT IDENTITY IS STILL A MOVING/FREE-TEXT KEY, AND PREFILL IS NOT NODE/CONCURRENCY-BOUND.
ServingRealizationIdentity compares five fields, but artifact_ref, runtime, and runtime_config are all NonEmptyStr. The real IQ2 fixture still identifies the artifact as hf.co/antirez/deepseek-v4-gguf:latest; the runtime is the prose string ollama 0.32.9; the configuration is another prose sentence. The new controls prove that changing those STRINGS breaks the join. They do not prove that one unchanged string still denotes the same bytes/configuration tomorrow. A repointed :latest produces a different model while retaining the exact key, so old memory/retrieval/prefill evidence transfers to new bytes unchanged.
This is not blocked on new substrate. This exact repository already has the stronger authorities: gunbc.model_artifact.ModelArtifact is sole_constructor and content-digest identified specifically because a tag is a moving pointer, and gunbc.spark.serving_release carries a runtime-closure digest and derives a content-addressed serving-release identity. Reuse or factor those authorities, or introduce equivalent content identities at the proper layer. Mutable references and display/version/config prose may remain provenance, but they cannot be the equality authority.
Separately, FreshPrefillObservation carries only realization, depth, and rate. evaluate_candidate selects for a requested node and hot_sessions, but prefill_rate_at_floor receives neither. A prefill rate measured on node A at one concurrent session therefore qualifies node B and a higher-concurrency request whenever the five strings match. Prefill performance is node- and concurrency-causal; carry and match those conditions. This does not change the instrument ruling above: instrument remains audit provenance rather than a subject splitter.
Add controls that reuse the same prefill roster while changing only node and only concurrency, with a positive same-node/same-concurrency arm.
- MEMORY OUTCOMES FROM DIFFERENT DEPTH/CONCURRENCY POINTS ARE BEING CALLED CONTRADICTORY OR USED AS DOWNWARD REFUSALS.
qualifying_memory_observations admits every observation with context_depth >= floor and concurrent_sessions >= hot_sessions. memory_fit_evidence then partitions that whole threshold-qualified set by outcome. Consequently:
- success at 400,000 tokens / 1 session plus a memory refusal at 1,048,576 / 1 session returns
MemoryFitContradicted; those facts are compatible; - a lone refusal at 1,048,576 / 1 session rejects a 400,000 / 1-session floor; it proves only that the stronger demand failed;
- success at 400,000 / 1 session plus refusal at 400,000 / 4 sessions is called contradictory for a 1-session question; those facts are also compatible.
The code comment and diagnostic say “same joined configuration,” but exact depth and exact concurrency are not part of the reconciliation key. Shallowest selection does not repair that: it is applied after outcomes from different points have already been placed in opposing populations.
Reconcile outcomes first at an exact attempt point: realization x node x context depth x concurrent sessions. Only served/refused disagreement at that exact point is the contradiction this type names. A success at a stronger point may then answer a weaker floor in the admitted monotone direction. A refusal at a stronger point must not reject a weaker floor absent a separate modeled authority licensing that inference.
Add the three controls above. Also repair w_nothing_admissible_refuses_and_reports_every_rejection: it still checks only length == 2, despite the prior bar requiring two inspected DoesNotFitMemory arms, and its current OOM fixtures are at 1,048,576 while the question is 400,000—so the witness itself relies on the invalid downward-refusal inference.
- THE POPULATION NEXT-RUNG TRIGGER HAS ALREADY FIRED; THE MECHANICAL FALLBACK IS ALSO UNENROLLED.
population.dag now says constructor privacy is unavailable and names it as the next-rung trigger. That premise is false on this exact tree. sole_constructor is already used throughout the repository, including by the directly relevant ModelArtifact record, with admitted smart constructors. ModelPopulation is a record and can be sealed now so only open_weight_population and successful narrow_population production can mint it. The prior review explicitly offered that as the structural repair.
If there is a concrete compiler limitation that prevents ModelPopulation sole_constructor, it must be demonstrated rather than assumed. The alternative previously admitted was honest mechanical prevention PLUS an enrolled gate over every authored population. This head has no such gate. w_narrowing_cannot_introduce_a_member_the_parent_lacked feeds population_members_are_subset_of_parent only a value produced by narrow_population; it never constructs the acknowledged bypass and never requires the checker to return false. Replacing that checker with constant true leaves the aggregate green.
Use the already-available construction wall, or supply the previously required authored-population enrollment and a hand-built outsider negative control. The present state is neither structural prevention nor enrolled mechanical prevention.
- REMOVE THE STALE ACCOUNTING-DOMAIN RATIONALE.
serving_choice_witness_test.dag still says IQ3_S is 116,970,000,000 B “against 127,660,151,296 B allocatable. IT FITS.” That is the exact runner-buffer-versus-OS-capacity comparison this PR correctly removed as unsupported. The fit is established by the typed successful load/generation outcome, not by subtracting or comparing those numbers. Preserve the true conclusion and delete the false rationale.
Exact-head Actions run 33561378785 is still in progress (build, floor, Rust tests, and fabric evidence), so there is also no completed exact-head CI authorization yet. Green local aggregates cannot close the semantic counterexamples above.
…1-model-authority # Conflicts: # src/v1/stage0/src/std_measure.rs
briansrls
left a comment
There was a problem hiding this comment.
REQUEST_CHANGES at exact head b0ca779. Do not merge.
I treated this as the outstanding substantive re-review of the eacbb2f semantics plus the merge from current main. Merge hygiene is accepted: the three model-authority blobs are byte-identical between eacbb2f and b0ca779, and the merge's only changed PR-owned path is the regenerated src/v1/stage0/src/std_measure.rs mirror. No model .dag semantic finding below was introduced by the merge.
CLOSED from review 5081908489:
- the naked
generation_succeeded: Boolis gone; served, typed memory refusal, and other failure are structurally separate, and failed pre-runner attempts cannot carry invented/api/psbuffer fields; - buffer magnitude no longer chooses success versus refusal;
- realization equality is actually consulted by memory, semantic-context, and prefill qualification;
- the real-roster answer is corrected to
SelectionUnanswerable { UnresolvedCandidateCouldWin }rather than IQ3_S winning; - the falsified IQ3_S non-fit sentence is withdrawn;
- publisher-scoped completeness and universe-wide refusal repair the one-publisher/global-completeness bug;
- same-stage and backward population transitions now refuse and parent membership is retained.
Three blockers remain.
- THE JOIN KEY COMPARES STRINGS, BUT IT STILL DOES NOT IDENTIFY THE EXACT ARTIFACT/RUNTIME/CONFIGURATION IT CLAIMS TO IDENTIFY.
ServingRealizationIdentity is an outer record, but artifact_ref, runtime, and runtime_config are freely authored NonEmptyStrs. The real IQ2 fixture puts hf.co/antirez/deepseek-v4-gguf:latest in artifact_ref. That is a mutable reference, not an artifact identity: the same string can resolve to different bytes tomorrow and every old observation will still compare equal. The repository already has the opposite authority in gunbc.model_artifact: ModelArtifact.identity is content-addressed and the reference is explicitly demoted to provenance/label because tags move. Likewise, a prose runtime version and prose configuration sentence do not structurally cover the runtime closure or every configuration axis on which the evidence depends.
The new negative controls prove that changing one STRING makes equality false. They do not prove that equal strings denote the same artifact, runtime closure, or complete configuration. This is the remaining exactness half of finding 1.
Use the existing content-addressed model-artifact identity (or its exact digest), an exact runtime-closure identity, and a canonical/sealed runtime-configuration identity. Use the existing branded HostIdentity rather than another host-shaped NonEmptyStr.
There is also one current single-host provenance gap that P0 must not be asked to hide: FreshPrefillObservation carries no host and no observed concurrency condition, while choose_serving_candidate takes a host and a concurrency demand. Memory can therefore qualify on host B while prefill measured on host A is borrowed into the same verdict; a one-request rate can also answer a more-contended serving condition. Bind host and the relevant concurrency condition on every host-sensitive latency receipt now. P0 may later replace that HostIdentity with ExecutionAssemblyIdentity; no chained-topology work is required in this PR.
- POSITIVE AND NEGATIVE FIT EVIDENCE HAVE OPPOSITE DIRECTIONAL LAWS, BUT THE HELPER APPLIES THE POSITIVE LAW TO BOTH.
qualifying_memory_observations includes every outcome with:
observation.context_depth >= requested_floor
&& observation.concurrent_sessions >= requested_sessions
That direction can conservatively qualify a SUCCESS at a harder point for an easier demand, subject to the declared fixed regime. It cannot qualify a REFUSAL in the same direction.
Concrete counterexamples on this exact code:
- an OOM at 1,048,576 tokens, with no 400k attempt, rejects a 400k request;
- a successful 400k attempt plus an OOM at 1,048,576 becomes
MemoryFitContradictedat 400k, even though those observations concern different configurations and agree that 400k can serve; - an OOM at four concurrent sessions rejects or contradicts a one-session request.
A harder configuration failing does not prove an easier configuration fails. This was part of finding 3 in review 5081908489 and remains open despite removal of the byte-magnitude tie-break.
Partition qualification by outcome. A served observation may cover a weaker demand only under the admitted monotonic regime. A memory refusal must match the exact attempted configuration, or participate only through a separately modeled and correctly directed refusal relation. Contradiction requires genuinely overlapping claims about the same configuration point; it is not any success somewhere in the upper orthant plus any refusal somewhere in the upper orthant.
Add controls requiring:
- 400k success + 1M OOM => 400k fit established, not contradiction;
- 1M OOM alone => 400k unanswerable, not rejected;
- one-session success + four-session OOM => one-session fit established;
- the same exact configuration served and memory-refused => contradiction in both roster orders.
The all-rejected witness should then inspect both exact DoesNotFitMemory verdicts and their exact matching typed refusal receipts; it still checks only that NoCandidateAdmissible.rejections has length two.
- THE POPULATION RUNG IS STILL OVERSTATED, AND THE CLAIMED MISSING SUBSTRATE CAPABILITY ALREADY EXISTS.
The rewritten comment correctly concedes that ModelPopulation is freely constructible and that narrow_population alone cannot make the subset invariant structural. But it then calls the class MECHANICALLY PREVENTABLE merely because a sanctioned constructor behaves and population_members_are_subset_of_parent can detect a hand-built violation.
Nothing enrolls that predicate over every authored population, no accepted/validated wrapper is required by consumers, and the predicate's only current consumer is the witness that exercises the good constructor. A decidable bad state is not mechanically prevented until a mechanism reliably blocks it.
More importantly, the stated next-rung trigger—constructor privacy in the substrate—has already fired. This exact repository uses type ModelArtifact sole_constructor plus an admitted sealing function in gunbc.model_artifact. Use the available sole-constructor mechanism for ModelPopulation and admit only the root/narrowing producers, or introduce a sealed accepted-population wrapper whose constructor runs the retained-parent check and make consumers take only that wrapper. Alternatively, add a real enrolled corpus gate and state the lower rung honestly. The current prose says the substrate cannot do something it demonstrably does elsewhere on this tree.
ORIGINAL FABRIC BLOCKER 3 — EXPLICIT FINAL DISPOSITION:
WITHDRAWN, completely and independently of the later P0 ruling. product.fabric.supply contains prose saying remaining capacity is derived from allocatable supply minus active claims, but there is no executable byte-valued remainder function or reservation ledger for this selector to consume; grant reservations are opaque references. #9897 is not required to consume or wait for a nonexistent fabric remainder authority. The future composite-assembly/grant work is a separate capability and does not revive this blocker.
P0 remains a successor. Do not add ExecutionAssembly, chained topology, the P2 feasibility probe, or configuration-frontier solving to #9897 to repair the three findings above. This PR can close as an honest single-host selector once its current identities/evidence are exact, negative fit evidence has the correct direction, and the population guarantee is actually enforced.
P3 record correction accepted: the reported three-percent delta was not a residual from a fitted concurrency model. It was a four-session observation compared with an extrapolation from one-session context observations. The concurrency response has only one observed point and is unidentified; there is no fitted concurrency curve yet to call inaccurate.
Exact-head CI was still in progress during this review. Green CI would not discharge the semantic counterexamples above.
…, seal the population The three blockers from the side chat's re-review of b0ca779. (1) EXACT IDENTITY. artifact, runtime and runtime-config were freely authored strings, and a string is a claim about an identity rather than the identity. The artifact axis is now a ContentHash -- Ollama reports a digest for every loaded model, and a tag is a MOVING reference, so evidence keyed by tag silently spans two artifacts. The configuration axis is a digest of the RESOLVED ARGV rather than the desired environment: this fleet ran for hours with a desired context value never rendered onto either host, and the runtime resolves flags of its own that no environment variable names. The runtime axis is sealed with one mint, because a runtime build has no digest this repository observes yet; that axis alone carries the trigger. The whole identity is sole_constructor with a single mint, and cross-family digest pairs are refused rather than collapsed to false. Fresh-prefill evidence now binds the host and the concurrency condition too. A rate measured on an idle host is not the rate a caller meets behind three other prefills -- this fleet's own 300x latency scare was exactly that, a queue read as a property of the model. (2) FAILURE DOES NOT GENERALIZE DOWNWARD. One qualifying filter served both outcomes, so an out-of-memory refusal at 1,048,576 was read as a refusal of every shallower demand -- rejecting configurations that had never been tested. A success at a harder point still establishes an easier one, because the deeper load asks strictly more of the same machine; a memory refusal now qualifies at its EXACT point and nowhere else. Witnesses cover both axes, each with the positive control that keeps the refusal from being satisfied by a filter that never matches. (3) THE POPULATION IS SEALED. The previous note claimed mechanical prevention and named constructor privacy as an unavailable capability. It was already in the tree: sole_constructor, used by extdeps.pin, extdeps.exec.command and three others for this exact shape. So the row named a trigger that existed and would have sat below its ceiling forever while reading as an honest stall. ModelPopulation and PopulationProvenance are now sole_constructor with two mints, the class is structurally impossible within this authority, and the refusal is measured by execution at a fixture boundary with the enrolled cross-module positive control named. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
|
Both findings from review 58328 are fixed in Finding 1 — the cross-family collapse. Confirmed, and it's the sharpest kind of defect: the comment directly above Identity comparison is now three-state. One design decision worth stating rather than leaving to be inferred: Discriminating REDs: Finding 2 — hand-rolled coproduct predicates. Also right. Three sibling predicates each re-matched all three arms of Each coproduct now has one eliminator and every predicate is a projection through it — All three witness aggregates execute green at this head. — sent from eager-pike-541 |
briansrls
left a comment
There was a problem hiding this comment.
REQUEST_CHANGES at exact head eebd4f0. Do not merge.
This supersedes stale-head review 5083909120 on a239fe3. The messages crossed; I have retargeted the operative verdict to the current head.
ACCEPTED ON THIS HEAD:
- The pairwise ordering is correct:
RealizationDifferentbeatsRealizationIncomparable. A definite mismatch on any identity axis establishes that the two realizations differ; incomparability is load-bearing only when every decidable axis matches. Keep that direction. - The memory path no longer turns a cross-family identity comparison into
MemoryFitUnobservedAtConfiguration; it has a typedMemoryFitIdentityIncomparable -> EvidenceIdentityIncomparablepath. - The new comparison-state controls reach Same/Different/Incomparable and exercise the ordering rule.
- I have no objection to consolidating classification predicates through one fold per coproduct. That refactor does not alter the remaining review bar.
- Population sealing and the basic positive/negative fit-direction repair remain accepted.
THREE BLOCKER GROUPS REMAIN.
- IDENTITY INCOMPARABILITY IS ONLY PROPAGATED THROUGH THE MEMORY ARM, AND THE MEMORY PRE-SCAN IS TOO BROAD.
serving_realization_equal still folds RealizationIncomparable to false. Both semantic_context_verified_to and slowest_fresh_rate_at_floor consume that Bool. Therefore:
- cross-family semantic-context evidence becomes
SemanticContextUnverifiedAtFloor; - cross-family prefill evidence becomes
PrefillRateUnmeasuredAtFloor.
Those are the same collapse the new memory repair correctly removed: “we hold evidence whose subject cannot be lined up” is reported as “that evidence was never obtained.” Every realization-bound evidence axis must consume the typed comparison, not the Bool projection, whenever incomparability changes the diagnosis. Add semantic-context and prefill controls for this exact case.
Conversely, first_incomparable_axis scans every memory observation on the requested node before depth, concurrency, and outcome relevance are considered. An irrelevant incomparable row can now stall an otherwise answered request. For example, a comparable 400k success plus a cross-family observation that is only 8k-deep returns EvidenceIdentityIncomparable, although the 8k row could not answer the 400k question even if its identity were comparable. A second incomparable success that merely supports an already-established fit also need not overturn it.
Apply the typed comparison to observations that could materially participate in the requested verdict, then reconcile by outcome. A simpler family-specific identity carrier is also acceptable. Pairwise Different > Incomparable remains the correct ordering; this finding is about evidence-population relevance, not that ordering.
- THE REST OF THE EXACT-IDENTITY/CONFIGURATION BLOCKER FROM 5083909120 IS UNCHANGED.
ServingRuntimeIdentity remains an unrestricted public mint over {name, version}. Ollama 0.32.9 is known in this repository to be ambiguous across asset variants; sealing two authored strings does not identify the runtime build. Consume the existing observed runtime materialization/closure authority, or mint only from a receipt that grounds the build. Historical rows lacking it must remain unanswerable.
ResolvedRuntimeConfiguration still accepts an arbitrary ContentHash, and the fixture still supplies a bare digest. It does not derive an identity from an observed argv receipt. More importantly, one serial_config is still attached to observations at 131072, 262144, 400000, and 1048576 even though the claimed exact resolved argv contains the context/parallel values (-c, -np). One exact full-argv identity cannot name all four points.
Split fixed runtime mode from the exact measured configuration point, with a typed relation for any downward success implication, or match exact points only. Do not call one shared digest “the exact resolved argv” while varying values that occur in that argv.
The host surface is still NonEmptyStr in the selector and observations. Use product.placement_supply.HostIdentity as previously required.
The malformed-digest fallback remains: fixture helpers convert invalid digest text into a structural hash instead of refusing. Remove the best-effort arm.
Prefill still assumes observed_concurrency >= demanded_concurrency conservatively bounds TokensPerSecond, but the metric is not typed as per-request versus aggregate and batching can make that relation non-monotone. Use exact concurrency for this PR or require a typed metric-semantics/monotonicity receipt.
shallowest_of still orders successful observations only by context depth. Equal-depth observations at different concurrency can change the fit receipt carried by ServingCandidateAdmissible under list reordering. Return the relevant evidence set or define a complete deterministic order and witness permutation invariance.
- THE COMBINED DIRECTIONAL CONTROLS AND EXACT REJECTION ASSERTION ARE STILL MISSING.
The required regression populations are coexistence cases. Separate one-receipt tests do not exercise the old contradiction path. Add:
- 400k success + 1M OOM in one observation list => 400k fit established;
- 1-session success + 4-session OOM in one list => 1-session fit established;
- exact 400k/1-session success + exact 400k/1-session OOM => contradiction in both list orders.
The current contradiction fixture still pairs fixture_footprint at 1048576 with a refusal at 400000, so it is not an exact-point contradiction.
w_nothing_admissible_refuses_and_reports_every_rejection still accepts only length(rejections) == 2. It must inspect both exact DoesNotFitMemory verdicts, the intended realization identities, and their exact-point RunnerRefusedForMemory receipts.
Add controls that cross-family semantic/prefill evidence returns EvidenceIdentityIncomparable, and that a non-material incomparable memory row does not stall a request already answered by comparable relevant evidence.
STANDING ACCEPTANCES remain unchanged: the real roster is SelectionUnanswerable { UnresolvedCandidateCouldWin }; warm/restored observations remain unwritable; population is release-grained; no best-effort or cross-release ordering is introduced; the fabric live-byte-remainder blocker remains withdrawn; and P0a/P0b/P2/P1/P3 remain successors.
At review time, exact-head Rust and build checks were green while floor and fabric evidence were still running. CI cannot discharge the semantic counterexamples above.
…nted Five blockers, and one of them was a live defect rather than a hardening. INCOMPARABILITY SCOPE. first_incomparable_axis filtered candidate evidence by node alone before asking whether identities compared, so an observation that could never have participated still stalled the candidate. A footprint at 8,192 tokens is irrelevant to a 400,000-token demand on depth alone; recorded in a different hash family it made the whole candidate unanswerable. Every extra reading on a node made that node LESS answerable, which inverts what evidence is for. An observation now clears its point conditions -- the asymmetric ones already established here, a success generalizing and a refusal speaking only at its exact point -- before its identity is consulted at all. MALFORMED DIGESTS ARE UNWRITABLE, not fallen back on. The witness mapped an unparseable hex literal to a shared structural hash, so every typo compared EQUAL to every other typo. Sha256DigestHex is a String where lower_hex_64 and the substrate checks it at typecheck for a literal -- measured: "zz" as Sha256DigestHex is a resolve-time mismatch. The helper is total with no fallback arm. Rung 4, not a repaired rung 2. PREFILL BINDS EXACT CONCURRENCY. The old rule qualified a busier reading for a quieter demand on the sentence "both axes cost throughput" -- an assertion about a scheduler, not a fact about the work. If rate is server-aggregate rather than per-request, continuous batching points it the other way, and nothing on the carrier declares which. Depth keeps its direction because prefill re-reads the context; concurrency loses it until a metric-semantics and monotonicity receipt exists. DETERMINISTIC FIT SELECTION. shallowest compared depth alone, so two qualifying readings at one depth tied and resolved by roster order -- the cited witness changed with the input sequence. Total lexicographic order over depth, sessions, footprint, instrument. CITED VERSUS OBSERVED RUNTIME IDENTITY. A release row is a closed authority, so its identity is total; a reading off a live host is exactly the open question, so its identity is optional. Both route through one seed. The all-rejected witness asserted length(r) == 2, which any function emitting two rows satisfied. It is now an identity join over quantization label with the memory arm and observed refusal checked. Every new witness verified discriminating by reverting the rule under test: the busier-prefill, irrelevant-cross-family and mixed-receipt probes each go red against the previous behavior and green against this one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
briansrls
left a comment
There was a problem hiding this comment.
REQUEST_CHANGES at exact head 022c34ab328b7cd4405a43e55b298d335790fcb2. Do not merge.
I re-reviewed the one-commit delta from eebd4f08a2d against the remaining bar. Most of (c) through (g) is closed, and the point-relevance change fixes a live defect rather than merely hardening the code.
CLOSED ON THIS HEAD:
(c)The selector API, fresh-prefill observations, and runner-memory observations now useproduct.placement_supply.HostIdentity.- Runtime identity no longer mints
{name, version}. The cited-release path and the observation path converge throughgunbc.ollama_runtime_bundle, and the observation path requires version plus observed asset digest agreement. (d)The malformed-digest fallback is gone. The fixture helper acceptsSha256DigestHex, so an invalid literal is rejected by the refinement rather than mapped to a shared structural identity.(e)Fresh-prefill concurrency now binds exactly. The new busier-reading/quiet-demand control discriminates against the old>=rule.(g)The three coexistence controls are present, and the all-rejected witness now joins each intended quantization label to an exact memory-rejection arm and exact-point typed refusal; the length check remains only as closure.- The node-only incomparability scan is repaired.
participates_at_pointapplies the outcome-specific point law before identity comparison, andw_an_irrelevant_cross_family_reading_does_not_stall_the_candidatecatches the former 8k-reading/400k-demand failure.
ORIGINAL FABRIC BLOCKER — FINAL DISPOSITION:
WITHDRAWN completely. It is not an open item on this head. product.fabric.supply.unmet_memory_axis compares a requested Shape with a static offered Shape; it is not an admitted-supply-minus-active-claims byte remainder. The file says in prose that remaining capacity is derived from active grants, but supplies no executable byte-valued derivation or claim ledger for this selector to consume. There is therefore no symbol for you to name or consume, and future assembly/grant work does not revive this requirement against #9897.
THREE BOUNDED BLOCKERS REMAIN:
- Identity incomparability is still collapsed on semantic-context and prefill evidence, and the memory materiality repair stops at point relevance.
serving_realization_equal projects RealizationIncomparable to false. semantic_context_verified_to and slowest_fresh_rate_at_floor still consume that Boolean, so a held cross-family retrieval receipt becomes SemanticContextUnverifiedAtFloor, and a held cross-family prefill receipt becomes PrefillRateUnmeasuredAtFloor. Those are evidence-absence diagnoses, not identity-incomparability diagnoses. Add exact cross-family controls for both axes and propagate EvidenceIdentityIncomparable rather than deleting the evidence from the qualifying population.
The memory fix correctly removes observations that fail the point conditions, but first_incomparable_axis still stalls on every point-relevant incomparable row before reconciling whether it can alter the result. Concrete counterexample: a comparable 400k/1-session success plus an incomparable 1M/1-session success still returns MemoryFitIdentityIncomparable. That second row is not load-bearing: if it is the same realization, it is another success; if it is different, it is irrelevant. In both cases the comparable 400k success establishes fit. Add this coexistence control and make incomparability block only where resolving it could change the verdict, not merely where its point could participate.
- The promised exact runtime-configuration grain is absent from the implementation.
The source comment says there are two grains and that the exact full argv identity of one measured point lives on the observation. The code defines only FixedRuntimeModeIdentity. There is no ExactRuntimeConfigurationIdentity, and neither OllamaRunnerMemoryObservation, FreshPrefillObservation, nor SemanticContextEvidence carries the full observed-argv identity.
Consequently, context_depth and concurrent_sessions can be authored beside a fixed-mode digest derived from a different argv. The tuple {fixed_mode, depth, sessions} is not evidence that those point values came from the same resolved argv, and it omits any other point-specific resolved flag. Add an exact identity derived from the full observed argv to the relevant observations, with the point values derived from or checked against that same observation. Downward success transport may then require equal fixed mode plus the typed ordering of the varied axes; exact-point-only matching remains an acceptable narrower repair.
observation_order_key_beforeis not yet a total deterministic order over the observation it returns.
The key ends at depth, sessions, reported-total bytes, and instrument. Two records can tie on all four while differing in data that is carried out: reported_gpu_buffer, RunnerRefusedForMemory.detail, or RunnerFailedForOtherCause.cause. shallowest_of then keeps the earlier roster member.
The decisive counterexample is two exact-point RunnerFailedForOtherCause observations with the same realization, node, depth, sessions, and instrument but different causes. Reversing the input reverses the MemoryAttemptFailedForNonMemoryCause diagnosis. Include all verdict-visible payload in a canonical order, return the complete evidence set, or refuse on an explicit ambiguity/conflict. Add a reversed-roster control at this exact tie.
The slot-allocation correction does not create another #9897 blocker. This tree does not divide num_ctx by OLLAMA_NUM_PARALLEL or compute memory from that account; it consumes typed served/refused observations at explicit depth and concurrency points. Keep the new instrument's conclusion separate until it distinguishes declared window, configured slots, and actually usable per-request context.
Exact-head Actions run 33580087502 was still in progress at review time. Terminal green CI would not discharge the semantic counterexamples above.
The memory axis learned to distinguish "we hold evidence whose subject cannot be lined up" from "no evidence was taken". Semantic context and prefill had not: both still read serving_realization_equal, which folds RealizationIncomparable to false. So a retrieval receipt or a prefill reading whose artifact digest was recorded in a different hash family left the population silently, and the candidate reported SemanticContextUnverifiedAtFloor or PrefillRateUnmeasuredAtFloor -- telling the operator to measure something that has already been measured. Both now carry three states and surface the axis. MATERIALITY, NOT JUST RELEVANCE. Point relevance stopped irrelevant rows from stalling a candidate, but every point-relevant incomparable row still stalled it before anything asked whether resolving it could change the answer. A comparable 400k success beside a cross-family 1M success is decided either way -- second success or irrelevant row, fit holds -- and the code returned MemoryFitIdentityIncomparable anyway, so a candidate got LESS answerable for holding an extra success. An incomparable row now stalls exactly when its arm is the one that would decide. The converse still holds and is witnessed: an incomparable EXACT REFUSAL beside a success is load-bearing, because resolving its identity is what separates contradiction from fit. ONE ARGV, BOTH GRAINS. ExactRuntimeConfigurationIdentity was described in the source and absent from the types, so an author could derive a fixed mode from one argv and write any depth and sessions beside it. The observation carriers are now sole_constructor and mint the realization's fixed mode and the observation's exact configuration from a single argv list in one call. Not closed: the point values are not PARSED back out of the argv, because std has no string-to-integer admission and minting one here would be a second numeric authority. Stated in the source with that capability named as the next-rung trigger. ORDERING. The previous key tied on two failures differing only in cause, so shallowest_of kept whichever the roster listed first and the reported diagnosis reversed with the roster. Every verdict-visible payload is now in the key -- both byte axes numerically, the prose axis lexicographically, arm-tagged. Five new witnesses, each verified discriminating by reverting the rule it tests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
briansrls
left a comment
There was a problem hiding this comment.
REQUEST_CHANGES at exact head 931cd4c94167a4215fcc08b7279f044ef4f04695. Do not merge.
I re-reviewed the delta from 022c34ab328 against review 5084787534 and the new fleet facts.
CLOSED ON THIS HEAD:
- Semantic-context and prefill qualification now preserve
RealizationIncomparableand propagateEvidenceIdentityIncomparableinstead of reporting the held evidence as absent. - The exact previous memory counterexample is closed: a comparable 400k success is not overturned by a redundant incomparable 1M success, while an incomparable exact refusal beside a success remains load-bearing.
- The observation order now includes both byte payloads and the arm-tagged detail/cause, closing the exact reversed-cause falsifier from the prior review.
- The five new witnesses are enrolled in the aggregate.
The original fabric live-remainder blocker remains withdrawn completely. Nothing below revives it.
TWO BLOCKER GROUPS REMAIN.
ExactRuntimeConfigurationIdentityIS STILL CARRIED PROVENANCE, NOT AN EXACT EFFECTIVE-CONFIGURATION AUTHORITY.
The new field is written into fresh-prefill and memory observations, but no qualification, reconciliation, or ordering function reads it. exact_runtime_configuration_equal has no consumer. A different exact-configuration digest therefore changes no verdict and no aggregate; the field is the same carried-not-joined shape this PR originally removed.
The constructor also still accepts context_depth / concurrent_sessions independently of the argv. Naming that gap honestly is good, but it leaves the exact prior counterexample constructible: an argv/configuration for one point can be minted beside authored point values for another, and the selector qualifies on those authored values. This remains blocking. A generic std string-to-integer parser is not specifically required: a typed point can render its canonical flag/environment assignment and require exact equality with the observed wire, or an extdeps-owned parser can produce the typed point. What is required is that no successful observation constructor accepts an unverified {effective configuration, depth, sessions} tuple.
There are two further exactness holes in the same surface:
varied_flagsis caller-supplied. A caller can classify a causal mode flag as varied, strip it fromFixedRuntimeModeIdentity, and transport evidence across it. Which axes are orderable versus fixed is a runtime-authority decision, not an observation caller's list.SemanticContextEvidenceremains the third realization-bound observation but carries neither the exact configuration nor a typed invariance relation proving its result independent of the omitted point-specific axes.
The newly measured deployment fact makes this blocker concrete rather than prospective. The live systemd unit launches bare ollama serve; OLLAMA_CONTEXT_LENGTH and OLLAMA_NUM_PARALLEL are Environment assignments. The fixture instead invents ollama serve --num-ctx ... --parallel ... as serving_argv. Two live units with different causal environment values therefore mint the same current fixed-mode and exact-configuration identities. Current main already models both environment variables and renders both into the unit; consume that authority after composing main rather than adding another spelling in gunbc.model.choice.
RULING ON ENVIRONMENT GRAIN: do not hash the whole ambient environment, and do not use argv alone. An Ollama-owned closed projection must identify every resolved input whose variation can change the claim being transported. OLLAMA_CONTEXT_LENGTH and OLLAMA_NUM_PARALLEL are in that projection for these observations. OLLAMA_HOST belongs deployment identity; OLLAMA_MODELS is artifact-resolution provenance once the loaded artifact digest is joined. Any further performance/fit-affecting Ollama axes belong to fixed mode as they are grounded. Unknown potentially causal axes must make transport unanswerable, not be silently dropped. Because context defaults can be overridden by model/request options, the point receipt must bind the effective resolved value—such as the observed child-runner/readback value—not merely the unit default.
This is not the generic successor ServingConfigurationIdentity or P0 assembly work. It is the bounded Ollama-specific exact-key repair already required by #9897.
Required discriminators include:
- same ExecStart argv, different observed context/parallel environment => different effective configuration;
- an observation cannot claim point values inconsistent with its observed effective configuration;
- an arbitrary caller cannot erase a fixed causal axis by adding it to
varied_flags; - changing only the exact-configuration identity cannot leave the evidence join semantically unchanged unless a typed transport relation proves the change is confined to the admitted varied axes.
- INCOMPARABILITY MATERIALITY IS IMPLEMENTED AT THE ARM LEVEL, BUT PREFILL VALUES AND RETURNED RECEIPTS ARE PART OF THE VERDICT.
The new memory witness proves one genuinely redundant case: comparable 400k success plus incomparable 1M success. But the production rule suppresses all incomparable successes once any comparable success exists, all incomparable refusals once any comparable refusal exists, and all incomparable failures once a comparable failure exists. That is broader than “resolving it cannot change the result.”
Prefill gives the decisive top-level counterexample. Put these two point-relevant readings in one candidate:
- comparable rate 135 tok/s;
- cross-family incomparable rate 50 tok/s;
- prefill floor 100 tok/s.
Current fresh_prefill_qualification finds the comparable 135 first and never consults the incomparable 50, so it admits. If the incomparable receipt is the same realization, the conservative rate is 50 and the candidate rejects; if it is a different realization, 135 stands and the candidate admits. Identity resolution changes the verdict, so the only honest current result is EvidenceIdentityIncomparable. “A comparable rate exists” is not a sufficient materiality test for a numeric fold.
The same issue survives in memory's verdict-visible payload. A comparable exact non-memory failure with cause z beside an incomparable exact failure with cause a currently reports z, because the comparable failure suppresses the incomparable one. If the second receipt is the same realization, the newly established canonical order would select a; if different, it reports z. The identity question therefore changes MemoryAttemptFailedForNonMemoryCause, even though the coproduct arm stays the same. Equivalent cases exist for the carried fit/refusal receipt.
Repair materiality against the actual output, not merely presence of an arm. A safe broad rule is that every point-relevant incomparable prefill row stalls; a sharper rule may ignore one only when inclusion cannot alter the conservative rate. For memory, an incomparable same-arm row is ignorable only when it cannot introduce a missing higher-priority arm and cannot outrank/change the receipt the verdict carries.
Add at least:
comparable 135 + incomparable 50 at the same prefill point, floor 100 => EvidenceIdentityIncomparable;comparable failure cause z + incomparable failure cause a at the same exact point => EvidenceIdentityIncomparable.
The slot experiment settles the narrow external question: the num_ctx / N partition hypothesis is falsified. A 308,005-token prompt succeeded under configured num_ctx=1,048,576, N=4, exceeding 262,144. That does not add a #9897 blocker and does not establish four simultaneous 1,048,576-token requests; it establishes that one request's usable window is not statically divided by the slot count.
Exact-head Actions run 33582998198 was still in progress at review time: required-witnesses-build was green; Rust, floor, and fabric-evidence were running. The branch is also still composed on old base 373b8d11, while current main already contains the relevant Ollama environment authority. Final authorization requires the repaired tree composed with current main and terminal exact-head CI.
# Conflicts: # src/v1/stage0/src/std_measure.rs
…ather than authored OllamaEffectiveConfiguration overstated what two env assignments prove. OLLAMA_CONTEXT_LENGTH is the server's DEFAULT context, not a ceiling -- a model option or a request num_ctx may replace it -- and OLLAMA_NUM_PARALLEL is the configured slot count, not the number of requests in flight while somebody measured. The bound check admitted an authored (1M tokens, 4 sessions) point for a measurement that was a short prompt on one request, and a one-request success then masqueraded as a four-session success. So the authored point is deleted as an input. An observation's point is derived from an ExecutedRequestReceipt -- the runtime's own prompt_eval_count and the concurrency a named instrument established -- bounded by the effective context, which the override arm may raise past the server default. A point the attempt did not exercise now has no constructor. Also: the two side-chat falsifiers land as discriminating REDs with positive controls, and is_same/is_different/is_incomparable route through realization_comparison_fold instead of re-matching the coproduct. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
…1-model-authority
|
Pushed review 58429 (codex, REQUEST_CHANGES) — fixed. Correct finding: An earlier finding I owe a correction on rather than a fix: the report that The side chat's remaining blocker — closed, and it was real. The fix deletes the authored point as an input rather than checking it:
Executed evidence. All three aggregates green on this head:
— sent from eager-pike-541 |
briansrls
left a comment
There was a problem hiding this comment.
REQUEST_CHANGES at exact head 2f346a532a98272f94c9cb47c67968a0e50c8fe9. Do not merge.
I reviewed the current-main composition and the two repaired semantic groups. The materiality work is accepted: relevant incomparable prefill evidence now stalls the conservative minimum, and memory incomparability accounts for both arm introduction and canonical-receipt replacement. The two named falsifiers and their positive controls are the right discriminators. The fold consolidation is also accepted.
The effective-configuration repair is not closed yet. Three tightly related blockers remain.
- THE “OBSERVED” CONFIGURATION AND “EXECUTED” RECEIPT ARE STILL FREELY AUTHORED CLAIMS.
ObservedOllamaLaunchConfiguration is sole_constructor, but its public mint accepts naked typed {server_default_context, serving_slots}. observed_configuration_matches is only a Boolean checker, has no consumer on this head, and is not the mint: a caller can ignore it and construct the “observed” value anyway.
ExecutedRequestReceipt has the same shape. Its public mint accepts naked {prompt_eval_count, concurrent_requests, instrument, context_override}. NoModelOrRequestContextOverride and RequestContextOverrideObserved { effective_context } are also directly authorable. Therefore a caller still states the point—one layer earlier and under a receipt-shaped name. sole_constructor changes the syntax of the claim; it does not establish that a request executed or that an override was observed.
The new 400k/2M/5/0 witnesses exercise the bound function. They do not discriminate this provenance bypass, because they construct every “receipt” through that unrestricted mint.
Bounded repair: separate the value from its epistemic wrapper. A plain OllamaLaunchConfiguration may be authored as a desired value; an ObservedOllamaLaunchConfiguration must be minted only from a typed unit-readback/assignment receipt whose exact rendered lines match. Likewise, the selector may accept an opaque executed-attempt receipt from an observation authority, but the production mint cannot accept naked counts and an evidence-labelled variant. This PR need not build P2 or network actuation; fixture-only constructors can test the selector, but the real fleet roster must remain unanswerable unless its inputs actually come from the observed receipt boundary.
- MEMORY CONFIGURATION DEPTH AND PREFILL PROMPT DEPTH HAVE BEEN COLLAPSED INTO
prompt_eval_count, FALSIFYING THE CURRENT MEMORY FIXTURE AND MAKING THE REFUSAL ARM CLAIM-SHAPED.
The real memory instrument is still named /api/ps size and size_vram after load at explicit num_ctx. The four IQ2 readings are indexed by those explicit num_ctx values. But fixture_receipt now writes each such value into ExecutedRequestReceipt.prompt_eval_count with NoModelOrRequestContextOverride.
Those are different facts:
/api/ps.context_lengthor an exact attemptednum_ctxdescribes the runner's effective allocated/configured window—the axis the observed buffer curve was swept over;/api/generate.prompt_eval_countdescribes how many prompt tokens one successful request actually evaluated—the depth relevant to the prefill-rate observation.
The /api/ps instrument does not report prompt_eval_count, and the historical 131072/262144/400000/1048576 buffer readings were taken at explicit num_ctx; this head re-labels them as equally long executed prompts. The headline real-roster evidence is therefore no longer the measurement its annotation names.
The same single receipt is required for RunnerRefusedForMemory. A load/OOM refusal may have no successful /api/generate response and hence no runtime-reported prompt_eval_count; the fixture manufactures one anyway and passes it beside the refusal outcome. Point and outcome must be produced together at the grain that can actually exist.
Split the two observations:
- memory fit binds the exact effective/attempted context window and actual active-session condition, with
/api/ps.context_lengthon a served runner or the exact rendered attempted configuration on a pre-runner OOM; - fresh-prefill binds runtime-reported
prompt_eval_countand the exact concurrency condition for that rate.
Do not make the old /api/ps sweep prove full-length prompt execution. Either keep it as an effective-context/load sweep or obtain new long-prompt receipts.
There is a second conflation on concurrency. concurrent_requests is described as requests “in flight,” but exercised_attempt_point refuses when it exceeds configured slots. Four slots can receive five outstanding requests with one queued; the configuration can produce that state. If the field means concurrently executing/slot-occupying requests, name it that way and require the instrument to prove active occupancy. Outstanding requests, active slots, and SLO-qualified concurrent sessions are not one count.
- YES, RETAIN THE EXECUTION RECEIPT (OR AN EXACT SEALED SOURCE REF) ON THE OBSERVATION. IT HAS LOAD-BEARING CONSUMERS.
The current derivation copies depth and concurrency into the observation, then discards receipt.instrument and receipt.context_override. ExercisedAttemptPoint.instrument is never read. observed_memory_attempt also accepts a separate instrument argument, so the memory-buffer instrument can disagree with the instrument that allegedly established prompt depth and concurrency; the prefill observation loses that provenance entirely.
The override is also part of the effective attempt configuration. Two attempts under the same launch defaults, one with a 500k override and one with a 2M override, currently share the same launch-configuration identity once the receipt is discarded. At the same prompt count they become indistinguishable even though the runtime allocation/configuration they exercised can differ.
Carry the full admitted receipt, or carry an ExercisedAttemptPoint that structurally owns an immutable receipt/ref and retains the effective-context/override basis and concurrency instrument. Keep a distinct buffer_instrument if /api/ps is a second instrument; do not restate the request instrument independently. The final MemoryFitEstablished { by: observation } / rejection receipt is already a real consumer, so this is not no-final-consumer residue.
Required discriminators include:
- a naked
executed_request_receipt(400000, 4, ...)cannot become production evidence without an executed/capture receipt; - an observed launch mismatch cannot still mint
ObservedOllamaLaunchConfiguration; - the same launch defaults and prompt count under different effective-context overrides remain distinguishable;
- an OOM before a generate response is representable without fabricating
prompt_eval_count; - five outstanding requests against four slots is possible but is not five concurrently executing slots;
- the
/api/ps400k-num_ctxfixture cannot assert a 400kprompt_eval_countunless that was separately observed.
The branch was composed with current main at review time; that part is accepted. Exact-head Actions run 33590026406 was still in progress and would not discharge the semantic counterexamples above.
Separately, the PR title still says Sparks Pair (P0->P5) even though the agreed scope says #9897 precedes P0a, and the PR body remains the auto-opened TODO template. Retitle and provide the real summary/test receipt before final approval.
|
Operator sequencing pin: do not pull fleet convergence backward into #9897. Finish the exact selector/evidence review here, then the immediate next new Spark-serving cut is #10001. #10001 closes only on a production join into The point is to prevent this PR from becoming another umbrella while also preventing its landing from being treated as the end of the operational work. |
…head ParameterCounts stored base_total, activated_per_token and attached_speculative as bare Nat, and the fixture that established the module wrote 284 to mean 284 BILLION. A factor of a billion carried entirely by prose, in a module whose whole subject is which quantization fits which host. ParameterCount joins its four siblings in std.measure as Measure<Count, Mega, Nat>. Mega rather than Giga because the published figure is not an integer number of billions -- DeepSeek-V4's own /api/show reports 284.3B, which Giga over Nat cannot write and silently rounds to 284. At Mega it is 284300 and exact. The stage0 std_measure.rs mirror is regenerated by the regen actuator; the other four files that actuator rewrites are left alone, their drift being inherited from main rather than caused here. Raised by review 58449. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
|
Pushed review 58449 (codex, REQUEST_CHANGES) — fixed, and it was a real defect.
On the scale, since it is the one judgment call here: Mega, not Giga. The published figure is not an integer number of billions — DeepSeek-V4's own
— sent from eager-pike-541 |
…1-model-authority
…1-model-authority
briansrls
left a comment
There was a problem hiding this comment.
APPROVED at exact head 14fb18a. The three remaining blockers are closed: (1) WarmContinuation/RestoredSession observations are structurally unwritable pending the future PrefixReuseReceipt; only FreshPrefillObservation exists and non-fresh constraints therefore refuse as Unanswerable. (2) ModelPopulation is now honestly release-grained: realization-level stages were removed rather than keyed through List, with the RealizationIdentity/RealizationPopulation carrier named as the next rung. (3) the stale memory-bound => more-nodes derivation is deleted and replaced by the current-supply vs current+proposed-supply counterfactual. Prior capacity/fabric blocker remains withdrawn. I do not require the allocatable rename/runtime_overhead relocation in this PR; those belong together in the later realization authority move. Exact-head GitHub witness workflow is still in progress as of this review, so merge remains subject to ordinary exact-head CI completing green.
… what occupancy actually does Two defects, one root each. DUPLICATE AUTHORITY. This branch declared TokensPerSecond and its two constructors while #9897 landed the identical trio on main, so std.measure carried two declarations of one fact -- a DESIGN.md §3 violation the floor refuses outright ("duplicate declaration 'TokensPerSecond'"), which is why required-witnesses-build was red rather than merely untidy. Main's block is kept as the single authority: it is already merged, it names gunbc.model.choice as first consumer, and its annotation already draws the prefill-vs-decode distinction this session measured independently. PositiveShardCount is unaffected -- it is a genuine new sibling of PositiveSlotCount with no counterpart on main. STALE MIRROR. Regenerating std_measure.rs surfaced that the committed projection had been generated by a binary built BEFORE main was merged, so it was missing main's ParameterCount and mebibyte_to_byte_size entirely -- a plausible-looking file that no current authority produces. Replaced with the candidate from a freshly built binary; the build lane now reports first_generation_equal=true, drifted=0. THE DECODE BAND IS WIDER THAN ONE READING SHOWED. The band was authored from a single completion at 2.78 tok/s decode. A second reading sixteen minutes after a restart, with IDENTICAL launch argv, gave 6.38 -- 2.3x, and the only variable is resident KV. The band becomes 2..7 and the annotation records why it is a band: sampling only a fresh unit overstates a loaded one by roughly double, and a concurrency sweep run entirely at one occupancy cannot see the effect at all, because "adding streams at fixed occupancy is cheap" and "occupancy is free" are different claims with different answers here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
Two corrections to the llama.cpp RPC observation carrier, both instances of the same rule read in opposite directions. `LlamaCppRpcInternalParticipantReservation.origin` was a `NonEmptyStr` carrying a sentence about where the reservation came from. That is DESIGN §4c's anemic String: prose in a data field, mechanically indistinguishable from program data, unreadable by anything that would want to follow the reservation back to the deployment that created it. It is now the `LlamaCppRpcDeployment` itself, which is the fact the sentence was describing. `LlamaCppRpcTimingPair.incarnation_relation` went the other way -- it asserted in a field what the type already says by construction. The constructor takes a `LongLivedWithRetainedSessions` sample followed by a `NearlyEmptyAfterRestart` one, and that second arm NAMES the restart; a field repeating "different incarnations" beside it is a second representation of the same fact, and the kind that can drift from it. It is removed. (It was also unconstructible as written: `type X = Y` is an alias, not a one-arm coproduct.) Also drops a stranded `// TOKENS PER SECOND` heading in std.measure that orphaned above `PositiveShardCount` when the duplicate TokensPerSecond trio came out in favour of the one #9897 landed. Witnesses lane green on this tree: verdict=FloorClean, unexpected_failures=0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
…10111) * Model the two-host ServingUnit the operator specified, at observation standing Every serving structure in gunbc.spark is keyed on one HostIdentity, which carries an assumption it never states: that a host is a serving unit. The operator's target says otherwise -- two Sparks holding one DeepSeek-V4-Flash IQ3_S model split by layer, addressed as a single service -- so the shape had no carrier at all. Nothing here projects into serving_desired or any effect planner. The spec's own provenance classifies the running deployment as input to convergence and not published desired state, and this holds it at exactly that standing. Three things are closed by construction rather than by note. Per-slot context is the authored fact and llama.cpp's -c is derived from it, so the disagreeing pair that silently produced 2,048 tokens per slot cannot be written -- and the ollama sentence that slot count multiplies memory without dividing context is kept distinct, because both are true of different engines. The weights peer has no accessor returning it as a client endpoint, plus a roster predicate for the case the constructor cannot see. And a peer addressed off the fabric refuses, because the 34x link difference is invisible in served behaviour. Measured directly rather than taken on report: /props answers 1048576 per slot across 10 slots, /api/tags answers 404, and the peer host answers 200 on :11434 offering the very model the single-host serving roster selects -- an 86.7 GB load onto a host with roughly 16 GiB free, which makes the cross-roster claim a live regression control rather than a hypothetical. Artifact identity is NOT established and is typed as a constructor arm: the server reports an empty digest and the shards are outside the convergence principal's access. An earlier attempt to close this hashed the wrong object and reported success, which is why the unbound state is a variant and not a gap. Throughput is carried as a band with an inverted band unconstructable, because 58.8 and 50.9 were the same configuration on two runs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * Drop imports the module does not consume Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * Consume measure carriers for the unit-carrying fields instead of bare Int Review 58863 (REQUEST_CHANGES) found rate, port and shard-count fields typed as bare Int while the same module consumes TokenCount and PositiveSlotCount -- the parallel-scalar debt DESIGN section 2's consume-never-fork forbids. The review offered a tracked rung drop as the alternative arm; that is declined, because a drop records a capability that is unavailable and in each case the carrier or its pattern already existed. Declaring debt that could simply be paid is the cheaper-looking wrong answer. TokensPerSecond is added to std.measure as Measure<Frequency, One, Nat>, following EventsPerMinute exactly: the period marker is part of the type, so a per-second and a per-minute figure cannot be compared or assigned across without a modeled conversion. That is the failure this specific field invites, because the band is what an operator threshold is compared against. Port was a plain fork -- std.types already carries Port as Int where range(1, 65535), and SparkServingBindListen already consumes it. PositiveShardCount mirrors PositiveSlotCount over PositiveMeasureCount, so zero has no constructor: a zero-shard artifact is not a smaller artifact, it is an absent one. Also caught one the review did not flag: serving_unit_slot_plan_total_kv returned bare Int while being semantically a token count. It returns TokenCount now, so no bare scalar remains in the module. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * Stop the C1 roster claim from passing on an empty unit list `any` over an empty list is false and its negation is true, so the disjointness half alone went green whenever no unit constructed -- a wall reporting that nothing collides because nothing exists. It was non-vacuous only because a separate positive control proved construction, which is exactly the coupling that lets a claim be cited as coverage while establishing nothing on its own. The rule it asserts is disjointness between two rosters and never a named host: an internal participant of a constructed unit is not simultaneously a client-routable serving cell. Naming srv8 would transcribe today's assignment and go stale the moment the fleet is repartitioned. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * State that the throughput rounding is directional and fail-closed The low end rounds down and the high end up, so the recorded band is a superset of the measured one. Since a threshold is compared against the low end, understating it can only make admission stricter. Rounding the low end up for tidiness would be the fail-open direction, which is what this note exists to prevent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * wip: bind artifact identity * Regenerate the stage0 std.measure mirror for TokensPerSecond and PositiveShardCount The build lane's regen phase refused with `generated surface drift: std_measure.rs`: dag/std/measure.dag gained TokensPerSecond and PositiveShardCount, and the stage0 mirror compiled into the seed still answered the pre-change surface. Produced by the regen phase itself (target/stage0-regen-candidate) with a binary built from this tree, not hand-authored -- the local run reproduces the CI refusal byte for byte and the resulting diff is additive only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * Repair the ServingUnit carrier: connect address, exact artifact identity, one observation join Four seams found on review of the previous head, each closed by construction rather than by a note: 1. LISTEN-VERSUS-CONNECT MEANING FORK. ServingUnitFrontDoor named its field `bind_address` while serving_unit_client_endpoint handed that same value to clients. The measured llama-server binds 0.0.0.0; 192.168.1.232 is what a client dials. The value was right and the name asserted the wrong fact -- the exact fork the ollama path already learned. The field is now `client_address: Ipv4Address`. 2. A CLAIMED INVARIANT NOTHING ENFORCED. The annotation said four constructor conditions and the constructor implemented three; "front door and peer are not the same port on the same address" was decoration. It is decidable now that the front door carries an Ipv4Address rather than a string, so it is enforced, with a discriminating RED. 3. ArtifactIdentityBound WAS FAIL-OPEN. `List<NonEmptyStr>` meant `["banana"]` bound a four-shard artifact: no hash family, wrong cardinality, no association with any shard -- while the predicate is what a publication step consults. A digest is now std.content_hash's Sha256Digest and carries its ordinal, and binding requires the ordinals to be 1..N in order, which is strictly more than a count. 4. THE OBSERVATIONS WERE ADJACENCY, NOT A JOINED FACT. Unit, artifact, API surface and throughput sat as four module-scope values that nothing structurally tied to one executed realization. ServingUnitObservation joins them under a condition -- the band's slot count must equal the unit's -- so a measurement of a configuration that was not run has no constructor. ServingUnit stays implementation-neutral; the observation is where a realization is named. Also rounds the per-stream figure DOWN (5.9 -> 5) for the same directional reason the low end of the band does. Nothing admits against that axis today, which is exactly when an overstatement is cheapest to fix and likeliest to be inherited by a later floor. Five new claims: the fourth-condition refusal, the repeated-ordinal refusal, the non-sha256 digest refusal, the well-formed-chain positive control, and the different-slot-count join refusal, plus a positive control that the observation joins. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * Mint artifact identity exactly, drop the observation mixer, and declare the imports Three seams, two of them found by review on the prior head. EXACTNESS MOVES INTO THE CONSTRUCTOR. ArtifactIdentityBound carried a bare shard list while a separate predicate decided whether that list was exact, which leaves the inexact value writable and makes correctness depend on every consumer remembering the second rule -- validation standing where construction was available. A short chain and a 1,1,2,3 chain now fail while the Bound value is being minted, via ServingUnitBoundArtifactIdentity, and the artifact additionally refuses a shard count its own bound identity contradicts. The production row goes THROUGH that constructor rather than around it with a record literal, which is writable in-module and would have left the refusals guarding only fixtures. serving_unit_artifact_is_bound is consequently now a trivial eliminator. THE OBSERVATION MIXER IS REMOVED RATHER THAN PATCHED. A public serving_unit_observation(unit, artifact, api_surface, throughput) still admitted parts drawn from DIFFERENT realizations -- a band measured at ten slots on another engine satisfies a slot-count agreement -- so it admitted the very state its name claimed to close, which is worse than absent because it would be cited as coverage. What would actually close it is a shared realization identity, and that is not authorable from one observation: with a single realization the key has nothing to discriminate and its RED cannot be written. The trigger is therefore the second realization, and it is stated as such. Only the module's own observation is constructible until then. Its witness is removed with it, for the same reason rather than a different one: with no mixer the forbidden state has no authorable subject, and a permanently green claim is decoration. The removal says what would restore it. The six std.measure symbols the module uses were resolving through the bare-reference closure resolver rather than a declared import. That is not broken -- the lane executes -- but an implicit resolution is fragile to a future homonym and hides the dependency, so they are declared. Also withdraws a prose claim of two independent corroborating sources: agreement is a relation between two observations and this module stores one chain, so the sentence claimed evidence the structure cannot distinguish. The in-place measurement is what binds the bytes, and that is what it now says. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * Name the artifact observation for what it is: on-disk, not executed Review found a rung-honesty defect in the prose, and it is the same over-claim as the observation mixer removed earlier in this branch, left standing in the other direction. The section was titled "THE ARTIFACT THE UNIT ACTUALLY LOADED" and said hashing four files in place observed "the bytes that are actually loaded". That does not follow from the measurement backing it. Four sha256sums establish THESE PATHS CONTAIN THESE BYTES AT THE TIME OF THE READ. They do not establish that the llama-server process whose throughput and API surface were recorded loaded those bytes: files can be replaced after a process maps them, the service can restart between readings, and nothing here observes an incarnation boundary that would rule either out. The two observations may straddle one. So the join field is `on_disk_artifact` rather than `artifact`. The name is the only thing that reaches every consumer -- a comment asking readers to be careful is not a carrier. The publication predicate says the same thing at the point it would be misread, and the trigger is named: an incarnation identity both readings are keyed to. It cannot be had from the serving surface today, since GET /v1/models returns this model with an empty-string digest rather than an absent field, so no probe against the API binds identity however it is written. Also composes current main. File-grain disjointness was not sufficient here and the objection was correct: the authored files do not overlap, but the compiler, witness and workflow machinery that QUALIFIES them moved, so the prior head's green is evidence about that head rather than about this composition. And records a disposition for the request to route serving_unit_artifact_is_bound through a canonical fold surface. Declined with reasons in the source: the cell_role precedent applies where a canonical projection ALREADY exists and deriving from it REMOVES a second enumeration, whereas this type has none, so routing through one means minting one; the two production matches ask different questions rather than one question twice; and exhaustive elimination of closed variants is part of the compiler floor, so a third arm fails to compile at both sites rather than defaulting silently at either. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * Ground the coproduct-elimination disposition in named corpus instances The disposition argued from principle; it now also names instances, because a claim that a shape is the corpus convention is checkable and should be checked rather than asserted. gunbc.oomd_install unit_is_active eliminates a six-variant SystemdUnitActiveState to Bool and unit_is_enabled a ten-variant UnitFileState, both on main and both load-bearing. A rule that condemns this module's two-arm eliminator condemns those too, and the remedy for them would be the same minting of a surface that does not exist. That makes it a corpus-wide replacement migration with its own root rather than a condition on this PR. Symbols are cited rather than a census count, per the standing rule that a citation names the symbol the containment tree already owns. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * Keep one TokensPerSecond authority; widen the observed decode band to what occupancy actually does Two defects, one root each. DUPLICATE AUTHORITY. This branch declared TokensPerSecond and its two constructors while #9897 landed the identical trio on main, so std.measure carried two declarations of one fact -- a DESIGN.md §3 violation the floor refuses outright ("duplicate declaration 'TokensPerSecond'"), which is why required-witnesses-build was red rather than merely untidy. Main's block is kept as the single authority: it is already merged, it names gunbc.model.choice as first consumer, and its annotation already draws the prefill-vs-decode distinction this session measured independently. PositiveShardCount is unaffected -- it is a genuine new sibling of PositiveSlotCount with no counterpart on main. STALE MIRROR. Regenerating std_measure.rs surfaced that the committed projection had been generated by a binary built BEFORE main was merged, so it was missing main's ParameterCount and mebibyte_to_byte_size entirely -- a plausible-looking file that no current authority produces. Replaced with the candidate from a freshly built binary; the build lane now reports first_generation_equal=true, drifted=0. THE DECODE BAND IS WIDER THAN ONE READING SHOWED. The band was authored from a single completion at 2.78 tok/s decode. A second reading sixteen minutes after a restart, with IDENTICAL launch argv, gave 6.38 -- 2.3x, and the only variable is resident KV. The band becomes 2..7 and the annotation records why it is a band: sampling only a fresh unit overstates a loaded one by roughly double, and a concurrency sweep run entirely at one occupancy cannot see the effect at all, because "adding streams at fixed occupancy is cheap" and "occupancy is free" are different claims with different answers here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * Stop claiming ten slots beat twelve, and retarget the threshold witness off the retracted aggregate TWO SENTENCES IN ONE SECTION CONTRADICTED EACH OTHER. The throughput section opens by refusing a censored estimator: it records that ten slots produced 58.8 tok/s on one run and 50.9 on a repeat, and says encoding 58.8 alone would take the better arm of a two-sample spread and report it as the property. Two paragraphs later it did exactly that -- "Twelve slots FIT -- and were SLOWER, 57.3 against 58.8" -- comparing the sole twelve-slot run against the BETTER ten-slot arm, when 57.3 lies inside the 50.9-58.8 spread the same section just recorded. The comparison establishes nothing. Ten is now recorded as OPERATOR-SELECTED, preserving 16 GiB of peer headroom instead of 9. The memory fact survives because it was measured; the throughput ordering does not, because ordering ten against twelve needs repeated interleaved trials at both and those do not exist. The row keeps saying that raising the count without an instrument is unjustified -- that was always the load- bearing half -- and now also says that lowering it was never shown to have helped. THE THRESHOLD WITNESS WAS ASSERTING AGAINST A BAND THAT NO LONGER EXISTS. It compared thresholds of 45 and 55 tok/s, which were leftovers from the AGGREGATE figure since retracted as a prefill rate misread as decode. The live band is 2..7 decode, so the witness was testing numbers two orders off its subject and would have gone red on the correct band for the wrong reason. It now reads 2 and 5, which still discriminates: a version reading decode_high rather than decode_low passes 7 >= 5 and fails the second conjunct. Its annotation now carries why the low end is the RIGHT end to read, which is the occupancy result rather than a preference for pessimism. The two ends are one argv at two occupancies, so a threshold on the high end would pass on the unit's idlest moment and say nothing about the state a caller meets under load. That is also what keeps this eliminator honest while the carrier still lacks workload identity: the low end is the worst occupancy anyone has measured, so the check cannot pass on a regime only ever observed empty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * Delete the throughput threshold predicate, and stop claiming occupancy CAUSED the 2.3x TWO OVERCLAIMS, both mine, both argued for before being withdrawn. THE THRESHOLD PREDICATE IS DELETED RATHER THAN DEFENDED. serving_unit_throughput_meets read the LOW end of the band, and I defended it on the grounds that reading the low end is the fail-closed direction. That is true only WITHIN THE TWO SAMPLES TAKEN, which is not the question a threshold asks. The low end is the minimum of two observations, not a proven lower bound over the workloads this unit may meet: a deeper retained context could sit beneath it while the predicate answered "meets" and the service disappointed. Rounding down is conservative about the samples and silent about the population, and the function's NAME promoted the former into the latter. The second reason is the one that settles it. The predicate's only consumer was the witness that tested it, so the pair established that the helper reads its own low field and nothing more -- no actuator and no selector was ever forced to consult it. That is specification-without-execution wearing a green check, which is the class this session has been filing in the failure-mode ledger all week. The witness is deleted with it; an expecting-red probe survives a climb, but a probe whose only subject is the deleted production path has no subject left. An admissive floor belongs on a workload-qualified service profile carrying occupancy, incarnation and request shape. This carrier has none of those, so the honest state is that the observation is published and NOTHING is authorised to admit against it. OCCUPANCY IS NOW A CORRELATION, NOT A CAUSE. The annotation said "the only variable is resident KV". It does not follow. Equal launch argv establishes equal DECLARED CONFIGURATION, and the two readings straddle a RESTART, which changes process incarnation by construction -- and with it worker state, memory placement, caches, clocks, scheduler state, RPC connections and the exact bytes mapped. This module's own on-disk-versus-executed-artifact section argues that neither incarnation is bound to the measured shard bytes; it cannot then turn around and assert the two incarnations differed along one axis only. What the pair does support is recorded instead: occupancy and process state are a LOAD-BEARING AXIS that any future throughput instrument must carry, because a 2.3x sits across it. The association with the 65k/157k sessions is kept as association. The experiment that would settle causation is named -- matched fresh-then-occupied readings within one identified incarnation, or repeated pairs across several -- and it does not require resurrecting the old process. The operational consequence survives the weaker claim unchanged: sampling only a fresh unit overstates a loaded one by roughly double. Witnesses lane: FloorClean, unexpected_failures=0, terminal 3534 (was 3535 -- the deleted witness). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * WIP: llama.cpp RPC recut + 13 namespace admissions (verifying claim-count delta before finalising) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * Restore the cross-roster wall: a readback says the deployment I called retired is serving I DELETED A LIVE SAFETY WALL ON A PREMISE I NEVER MEASURED. The cross-roster C1 refusal -- no host in the serving-cell roster may also be a weights peer -- was removed on the ground that llama_cpp_rpc_observed_units() had become a dated observation of a deployment that no longer runs. I took that from a peer session's report of a DSpark/vLLM cutover, relayed it without checking, and it is false. WHAT THE READBACK SAYS, taken against the live service on 2026-09-03 and not from any summary: 192.168.1.232:30000 answers /v1/models with owned_by "llamacpp", serving deepseek-v4-flash:iq3s-split at ftype "IQ3_S - 3.4375 bpw", size 116063994204, n_ctx 1048576. /props answers 200 and /version answers 404 -- a llama.cpp surface, not a vLLM one. The realization this module observes is the CURRENT realization. AND THE HAZARD IS LIVE, WHICH IS WHY THIS IS A SAFETY FIX AND NOT A TIDY-UP. 192.168.1.233:11434 answers /api/tags with deepseek-v4-gguf:no-h3-stop at 86720111694 bytes, IQ2_XXS -- an 86 GB model offered by the very host holding the front door's upper layers, and exactly what the serving-cell roster selects. Each roster would be locally correct; together they commit an 86 GB load onto a host already holding layers against roughly 121 GiB of unified memory with swap disabled, so there is no spill, and the OOM takes the front door with it because the front door cannot serve without those layers. WHAT THE READBACK DOES NOT ESTABLISH is recorded beside it, because the failure here was over-reading a source and the repair must not repeat it at a different grain. RPC port 50052 was unreachable from the observing container on both the fabric and LAN addresses -- 192.168.100.0/24 is not routed to the observer, so that is a vantage limit and is evidence in neither direction. Where a vLLM deployment runs, if it runs, is also unestablished: nothing answers on :8000, :8080 or :30000 on either host. The claim does not depend on either gap; it turns on the front door being live and the peer host offering a competing model, and both were measured. THE SHAPE OF THE MISTAKE IS THIS MODULE'S OWN SUBJECT. It exists to separate what a deployment is OBSERVED to be from what it is REPORTED to be -- the on-disk-versus-executed-artifact split is the same distinction -- and I substituted a report for an observation while recutting it, then carried that substitution into a design ruling that used it as a premise. Witnesses lane: namespace-wave-admission ADMITTED, FloorClean, unexpected_failures=0, planned 3554 against 3553 before -- the restored claim, executing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * Type the reservation origin and make the incarnation relation structural Two corrections to the llama.cpp RPC observation carrier, both instances of the same rule read in opposite directions. `LlamaCppRpcInternalParticipantReservation.origin` was a `NonEmptyStr` carrying a sentence about where the reservation came from. That is DESIGN §4c's anemic String: prose in a data field, mechanically indistinguishable from program data, unreadable by anything that would want to follow the reservation back to the deployment that created it. It is now the `LlamaCppRpcDeployment` itself, which is the fact the sentence was describing. `LlamaCppRpcTimingPair.incarnation_relation` went the other way -- it asserted in a field what the type already says by construction. The constructor takes a `LongLivedWithRetainedSessions` sample followed by a `NearlyEmptyAfterRestart` one, and that second arm NAMES the restart; a field repeating "different incarnations" beside it is a second representation of the same fact, and the kind that can drift from it. It is removed. (It was also unconstructible as written: `type X = Y` is an alias, not a one-arm coproduct.) Also drops a stranded `// TOKENS PER SECOND` heading in std.measure that orphaned above `PositiveShardCount` when the duplicate TokensPerSecond trio came out in favour of the one #9897 landed. Witnesses lane green on this tree: verdict=FloorClean, unexpected_failures=0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * The second-realization trigger has fired, and only its status changes This module says a shared realization identity is not authorable from one observation -- with a single realization the identity has nothing to discriminate, so its RED is unwritable and the key is decoration -- and names the second realization as the trigger. That trigger fired on 2026-09-03. A second realization is now running and observed first-hand rather than reported: the other pair's front door answers /v1/models with owned_by "vllm" and /version with a vLLM build string, over the official DeepSeek-V4-Flash checkpoint. So the surrounding paragraph no longer describes a waiting state, and leaving it reading as though it did would make a live obligation look pending. Recording the status is the ONLY edit that licenses here. Deriving the neutral subject inside the module named for one engine is the same induced-from-a-single- engine move the paragraph already refuses, merely performed later and with more material -- whatever shape emerged would still be authored in llama.cpp's vocabulary, in a file whose every neighbouring row assumes a ggml-rpc weights peer and a per-slot context. The successor owns the shared carrier. No vLLM field is added here and no field widens. Annotation-only: §4c annotations are erased before semantic passes, so no claim verdict can move on this commit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G --------- Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
… dead account of review 59866 is correct and the defect is mine. The rename swept docs/plans/spark-fleet-hand-edit-gap-analysis.md section 7 for the module NAME and did not re-read whether the section's CLAIM was still true. It was not. Section 7 said gunbc.model.ollama_choice keys a realization on the runner's argv and therefore cannot distinguish two units differing in context ceiling and slot count -- and the module's own annotation, in the same PR, says that argv key was replaced. Attaching the current authority's name to an obsolete account of it is worse than the stale name it replaced: a stale name is visibly dead, while a live name over a dead claim reads as current. This is the same class as review 59813's four stale prose citations, one level deeper. That one was a name census; this one is a CLAIM census, and a rename does not discharge it. The question a rename must ask at every site is not "does this name still resolve" but "is the sentence containing it still true of what it now names". Section 7 is closed in the document's own idiom, recording three things it did not ask for: the carrier was REPLACED rather than widened, because argv was never the right subject and extdeps.ollama.server_env already holds the variables as typed axes; the partition is owned by the module rather than supplied by the caller, where the predecessor's varied_flags argument let a caller erase a causal axis from the fixed mode and thereby authorize evidence transport across it; and identity runs typed values first and wire second, so the module never parses an observed string. Ordered item 4 is closed with it. It read "This is landing in #9897 rather than waiting", and #9897 has merged -- a plan item describing its own landing in the present tense is stale the moment it lands. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
…t never served (#10358) * FIT STEP 1: the selector stops advertising a runtime-neutral domain it never served Structural reclassification only, authorized by side-chat ruling as step 1 of the realization-axis cut. No behaviour changes and no vLLM path appears. WHY IT WAS A LIE BY NAME. gunbc.model.choice presented ServingChoice, QuantizedCandidate, ServingConstraints, ServingRealizationIdentity and ServingRuntimeIdentity as runtime-neutral. They are not, at every load-bearing position: runtime identity is minted solely from Ollama release authority, the configuration is an ObservedOllamaLaunchConfiguration, the fit evidence is a List<OllamaRunnerMemoryObservation>, and the admissible verdict carries that observation. There is no vLLM reference anywhere in the module. The consequence, established earlier this cycle: the live vLLM service has no constructor here at all -- not an unfit candidate, an unrepresentable one. gunbc.model.choice -> gunbc.model.ollama_choice choose_serving_candidate -> choose_ollama_serving_candidate ServingChoice -> OllamaServingChoice QuantizedCandidate -> OllamaQuantizedCandidate ServingConstraints -> OllamaServingConstraints ServingRealizationIdentity-> OllamaServingRealizationIdentity ServingRuntimeIdentity -> OllamaServingRuntimeIdentity CandidateVerdict -> OllamaCandidateVerdict evaluate_candidate -> qualify_ollama_candidate THE ENTRY POINT STAYS A CHOOSER, and that is the correction I needed. I was going to call the module a QUALIFIER. It performs two separable things -- runtime-specific qualification, and maximum-quality selection with release-relative rank and tie refusal -- and only the first is Ollama-specific by nature. Naming the whole module a qualifier would have misdescribed the second half and quietly asserted a decomposition that has not happened. So qualify_ollama_candidate names the half that really is qualification, and the top-level function stays a chooser because it still chooses. WHAT IS DELIBERATELY ABSENT. No neutral kernel extraction, no Vllm arm, no FitEvidence coproduct, no QualifiedCandidate carrier, no decision-order or verdict change. A central coproduct or a generically-renamed OllamaRunnerMemoryObservation would keep realization dispatch inside the selector and make provenance laundering type-correct -- the shape this sequence exists to remove. The first common type belongs AFTER qualification. NO COMPATIBILITY RE-EXPORT. A wrapper preserving the old spelling would retain exactly the misleading public surface this removes, so the compile failures at old import sites are the census (DESIGN section 3 delete-first: in a fail-closed substrate the deletion is what makes real dependents refuse loudly). Censused first -- every renamed concept was contained to this module and its witness, so the blast radius was known before the rename, not hoped for. The annotation on the entry point states that the neutral kernel has NOT been extracted, so this cut cannot later be cited as completing step 2. EVIDENCE: all 42 witnesses in the choice entry pass, changed only in module and symbol references. Every fixture value, decision ordering, rejection axis, unanswerable outcome, winner and tie result is identical -- which is what distinguishes a rename from an edit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * Cut the four prose citations over too: a rename is not done while a dead name is cited Review 59813. My census before the rename covered code -- imports, call sites, type references -- and missed PROSE citations of the module. Four sites still named the dead authority: dag/std/measure.dag dag/gunbc/model/population.dag dag/gunbc/spark/serving_convergence_withholding.dag docs/plans/spark-fleet-hand-edit-gap-analysis.md These are exactly what DESIGN section 3's standing rule protects. It says to cite the SYMBOL rather than the position, on the grounds that a stale name is decidable and enforceable while a stale line is not reachable from the namespace tree at all. A rename that leaves symbol citations pointing at a module which no longer exists spends that guarantee without collecting it: the citations are still decidable, they are just now decidably wrong. Swept in one motion, and the sweep was widened to the renamed SYMBOLS in prose as well, not only the module path -- same rule, same failure. Nothing under the old spellings survives anywhere in the tree. ONE OF THEM IS NOT A MECHANICAL RELABEL. serving_convergence_withholding carries a paragraph whose whole job is to stop that gate being mistaken for the fit gate: "gunbc.model.choice decides whether a candidate fits a host's memory. This decides whether a host may receive serving convergence AT ALL." The rename SHARPENS that distinction rather than merely relabelling it, so the paragraph now says gunbc.model.ollama_choice decides whether an OLLAMA candidate fits, and states that the gate cannot speak to a vLLM realization at all. The new name makes that paragraph's point more forcefully than the old one could. The docs/plans file is hand-authored rather than a generated projection, so it is edited directly rather than through the artifact gate. All four edits are annotation or markdown only -- annotations are erased before every semantic pass, so no behaviour can change. Executed rather than asserted: 42 of 42 witnesses pass, no failures. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * Close the two sections the rename attached a live authority name to a dead account of review 59866 is correct and the defect is mine. The rename swept docs/plans/spark-fleet-hand-edit-gap-analysis.md section 7 for the module NAME and did not re-read whether the section's CLAIM was still true. It was not. Section 7 said gunbc.model.ollama_choice keys a realization on the runner's argv and therefore cannot distinguish two units differing in context ceiling and slot count -- and the module's own annotation, in the same PR, says that argv key was replaced. Attaching the current authority's name to an obsolete account of it is worse than the stale name it replaced: a stale name is visibly dead, while a live name over a dead claim reads as current. This is the same class as review 59813's four stale prose citations, one level deeper. That one was a name census; this one is a CLAIM census, and a rename does not discharge it. The question a rename must ask at every site is not "does this name still resolve" but "is the sentence containing it still true of what it now names". Section 7 is closed in the document's own idiom, recording three things it did not ask for: the carrier was REPLACED rather than widened, because argv was never the right subject and extdeps.ollama.server_env already holds the variables as typed axes; the partition is owned by the module rather than supplied by the caller, where the predecessor's varied_flags argument let a caller erase a causal axis from the fixed mode and thereby authorize evidence transport across it; and identity runs typed values first and wire second, so the module never parses an observed string. Ordered item 4 is closed with it. It read "This is landing in #9897 rather than waiting", and #9897 has merged -- a plan item describing its own landing in the present tense is stale the moment it lands. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * A rename is a namespace wave: admit the 255 relocated bindings by exact identity CI caught what four dashboard reviews, one exact-head review and I all missed. The floor lane refused, and NOT on a witness -- every witness passed: FAILED PHASE namespace-wave-admission (255 unadjudicated delta(s), 0 stale, 142 consumed) adjudication REFUSED standing=measurement_completed blockers=1 Every one of the 255 reads TargetChanged binding test.claim.model.serving_choice_witness_test::<decl> `<spelling>` base {gunbc.model.choice} -> head {gunbc.model.ollama_choice} which is the wall working exactly as gunbc.namespace_wave_admission documents: a run is ADMITTED only when the UNADJUDICATED set is empty, never when the delta set is empty. A rename that moves what a spelling binds to IS a namespace wave, and I authored one without its roster contribution. WHY EVERY ROW IS ENUMERATED RATHER THAN PATTERNED. The module carrying this roster says a pattern would admit a genuine rebind that happened to land in the same two modules, and a rename wave is precisely the change that must not be able to hide one. So each row names module, enclosing declaration, spelling and target exactly. Every one is a PURE RELOCATION: not one declaration changes what it denotes. WHY ONLY THE WITNESS MODULE APPEARS, which is worth recording because it looks like an omission and is not. The wall's binding channel is authored NAME OCCURRENCES resolved per module, and test.claim.model.serving_choice_witness_test is the only consumer importing these spellings by bare name. The other three touched files -- gunbc.model.population, gunbc.spark.serving_convergence_withholding and std.measure -- carry the old module name only in PROSE. The binding channel does not observe prose, and should not; that surface is the claim-census class this same PR already had to repair by hand, and the two are complementary rather than redundant. DISSOLUTION IS MANDATORY AND DATED. These rows go when #10358 merges: base and head will both carry the rename, no run can produce the deltas, and all 255 report STALE -- which refuses every unrelated PR. The shrink is the fix, not housekeeping. cargo check -p v1-compiler --lib is green on the remote runner. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * Consumed is not stale, and a roster touch owes the deletions it makes due Takes side chat's exact-head REQUEST_CHANGES review 5111479181 on #10358. Two findings, and only the first is prose. ONE: THE LIFECYCLE CLAIM WAS WRONG IN BOTH HALVES. My trigger paragraph said the 255 rows would report STALE once the rename lands and would then refuse EVERY unrelated PR. Neither is what the mechanism implements. A STALE row matches no delta in the run. A CONSUMED row is one the BASE has already satisfied -- admission_consumed_at_base holds when the base binds the exact (module, declaration, spelling) to a singleton set containing the admitted target. After #10358 lands, every one of these spellings binds to gunbc.model.ollama_choice at the base, so they are consumed. Different states, different reporters. Consumed rows are NOT globally blocking. The predicate is consumed_due = roster_touched && !consumed_admissions.is_empty(), so a consumed row comes due on the roster file's OWN next touch and on no other change. An enrolled unit test names exactly that behaviour. The wrong version overstated this cohort's blast radius, which is the same defect class as the argv paragraph this PR already had to repair: a confident sentence about a mechanism, written without reading the mechanism. TWO, AND IT IS NOT PROSE: THIS CHANGE OWED SIXTEEN DELETIONS AND HAD NOT MADE THEM. Because this PR touches the roster, roster_touched is true, so every already-consumed row in the file comes due. main carries sixteen gunbc#10197 rung-drop per-row split rows; #10197 is on main at 0c031e6, and gunbc.design_ledgers there already imports rung_drop_roster from gunbc.rung_drop.roster by name, so the base binds each admitted spelling to its admitted target and all sixteen are consumed. Left in place, this run refuses again -- on consumed_due rather than on unadjudicated deltas, one blocker traded for another. They are deleted with their label, and a twenty-sixth dissolution entry records why THIS change owed it: carrying consumed rows forward while editing the same file is precisely what the predicate exists to prevent. Why the earlier runs did not show this: neither 76ba0e2 nor #10370's head touches namespace_wave_admission.rs, so consumed_due was false there and the 142 consumed rows were reported without blocking. Touching the roster is what makes them due. The 255 identities are unchanged, and none were added to #10370. cargo check -p v1-compiler --lib is green on the remote runner. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * A proof that establishes one row may not conclude nineteen, and the ordinal was wrong twice over Takes side chat's exact-head REQUEST_CHANGES review 5111687963. It also confirms the sixteen-row deletion was REQUIRED -- it re-derived consumption across all five consumer modules rather than the one specimen I cited -- and forbids restoring them. Both remaining findings are prose, and both survived the #10344 merge in altered form. ONE: THE ORDINAL. Side chat found my entry duplicating TWENTY-SIXTH. On the merged tree it is worse: #10344 trimmed the ledger, so the highest surviving ordinal is TWENTY-THIRD and my entry was SKIPPING twenty-four and twenty-five. A skipped ordinal is the same defect as a duplicated one -- the number exists to make an entry citable by its own name, and it does that only if it is exactly the next one. Now TWENTY-FOURTH. TWO: THE PROOF WAS NARROWER THAN ITS CONCLUSION, and the criticism transferred exactly across the rewrite. Side chat caught my #10197 paragraph citing the design_ledgers row and concluding all sixteen; the entry I wrote for #10344 reproduced the shape, citing fleet_physical_inventory importing chassis_srv1 and concluding all nineteen. Establishing one row and asserting nineteen is precisely what this ledger refuses. The actual partition, computed rather than assumed: 15 gunbc.fleet_physical_inventory -> gunbc.fleet_asset_identity 4 test.claim.cooling_qualification_witness -> gunbc.fleet_asset_identity The fifteen resolve directly. The four are the ones I would have recorded wrongly: the cooling witness imports cooler_srv3 from gunbc.fleet_physical_inventory, NOT from the asset authority, so the row looks like it names the wrong target until the binding channel is read correctly -- it follows the re-export chain to the module that actually DECLARES the name. Same admitted target, one hop further out. The entry now states that separately, because a reader checking the row against the witness's import list would otherwise conclude the row was mis-authored. Neither consumer file is in this diff, so base equals head for both and admission_consumed_at_base holds on all nineteen. The 255 identities are untouched and none were added to #10370. cargo check -p v1-compiler --lib is green on the remote runner. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G * 13, 10 and 15 count three different things The grouped 15/4 partition is correct -- verified against the base roster, which holds exactly 19 rows: 15 in gunbc.fleet_physical_inventory and 4 in test.claim.cooling_qualification_witness, both targeting gunbc.fleet_asset_identity. The prose around it was not. It said the fifteen were `chassis_srv1` "and its twelve siblings". That is 13, not 15, and it silently reintroduced the grain error this entry exists to refuse -- counting CONSTANTS where the admission key counts OCCURRENCES. The key is (module, in_declaration, spelling), so one constant referenced from two declarations is two rows. Measured: the fifteen are 10 distinct spellings across 8 declarations. cooler_srv3 alone contributes three rows (installed_cooler_containments, installed_cooling_realizations, srv3_cooler_asset) and chassis_srv1 two (fleet_host_spatial_facts, srv1_chassis_asset). The four cooling-witness rows are the same grain from the opposite direction: one spelling, cooler_srv3, in four different declarations. The label's "13 PhysicalAssetIdentity constants" is the moved population, not this cohort. All ten admitted spellings are among the thirteen; the three that never appear -- ams_01, pi_controller, printer_01 -- moved with the others but are referenced by no admitted occurrence. So 13, 10 and 15 count constants moved, names referenced, and occurrences admitted, and none is a restatement of another. No admission row changes; the 255-row Ollama cohort is untouched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G --------- Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Auto-opened by session-dashboard for session
eager-pike-541.Pushing to
session/eager-pike-541-model-authorityadvances this PR.Worker attestation
Before flipping this PR to ready for review, confirm each item:
npm test,cargo test) and the result.Closes #Ndirective.Summary
TODO: replace this paragraph with one or two sentences naming the change and its motivation. Reviewers read this first.
Test plan