Repository navigation
The serving cell's context value was a default misnamed as a limit, its concurrency axis was unmodeled, and the cost model for the trade was false - #9960
Conversation
…han a limit, and file the drop Three fleet-serving facts that were wrong, missing, or misnamed. MISNAMED. extdeps.ollama.server_env called OLLAMA_CONTEXT_LENGTH "the context window the SERVER allocates" and "the CONFIGURED limit". Upstream 0.32.9 resolves options in layers -- environment, then model, then request -- and an explicit num_ctx at either later layer replaces it. So it never limited anything; it is the value used when no override applies. One name carried two materially different contracts, and a deployment reading it as a ceiling would believe it had bounded something any caller can raise. Renamed through the carriers to default_context. MISSING. OLLAMA_NUM_PARALLEL was not in the closed variable set and not rendered, so the hosts inherited whatever the runtime chose -- one slot, which SERIALIZES. On 2026-09-01 a ~61,000-token prefill held one host for roughly two hundred seconds while every other caller queued, and the queue was read as a latency regression in the model. The axis is now modeled and the desired value is 4. WRONG. The P1b note said the two axes trade "because the slot count divides the context window". Measured on one host, one model, changing only that variable: one slot reported context_length 1048576 and 88,865,253,620 bytes; two slots reported context_length 1048576 and 90,543,761,652. Each slot gets its own full window. The count MULTIPLIES memory; it does not DIVIDE context. Deleted rather than softened. The desired context becomes 1,048,576, and its evidence sentence is replaced. The old one cited an /api/ps reading as proof of a configured value -- but /api/ps reports what a loaded runner holds, and both hosts carried no OLLAMA_CONTEXT_LENGTH line at all, so the number being confirmed was the runtime's own default. A declaration cited its own absence as confirmation. AND THE DROP IS FILED, one row for both values because they are one defect in two spellings: measured on one node, one artifact, one runtime, one concurrency condition, applied to two hosts and every realization the cell loads. The trigger names the capability rather than an artifact -- an observation table alone does not discharge it, because while a global default exists an unmeasured realization still inherits these numbers automatically. An unobserved realization must refuse, not inherit.
…criminating witness, and the figures give way to their instrument Three findings from review 58289, all valid. A BARE Nat CARRIED THE SLOT COUNT into the substrate, when std.measure is the authority for domain quantities. It is now PositiveSlotCount, built over PositiveMeasureCount the way PositiveCelsiusDelta is -- which also answers the second half of that finding, that the old carrier could represent a zero-slot configuration. Zero now has no constructor rather than a validator: a serving cell that serves nothing is not a quieter configuration, it is an absent one, and a caller who wants that removes the cell. THE RENDER HAD NO DISCRIMINATING TEST FOR THE NEW DIRECTIVE. Every existing arm varied the window; the slot count was always the desired value, so OLLAMA_NUM_PARALLEL could have stopped being rendered with the whole suite still green. That is the same condition the window arms were written to prevent, one axis late. Three arms added that hold the window fixed and vary only the slots -- bytes, content identity, and the causal link to the declared value -- each stated in both directions so a renderer ignoring its argument fails the equality half. THE ANNOTATIONS TRANSCRIBED RUNNER FIGURES rather than naming a producer that re-derives them, which DESIGN section 6 forbids for the reason this file already demonstrates: the sentence being replaced cited an /api/ps reading as proof of a value that had never been rendered onto either host, and nothing connected the number to the thing that owned it. Every byte figure is removed. The instrument is named instead -- gunbc.spark.serving_observe spark_serving_observe_ci_wet, whose residency readback in gunbc.spark.serving_readback_parse already decodes the runner's size_vram. The arguments survive as comparisons rather than as copied numbers: the same context_length at one slot and at two with a larger buffer at two still says the count multiplies memory and does not divide context. stage0 mirrors PositiveSlotCount alongside PositiveCelsiusDelta; cargo check and clippy clean. Spark witnesses re-run: unit render including the three new arms, release, observation.
… ones required-witnesses-build failed on 670873f. The cause is not the modeling: it is that I hand-wrote src/v1/stage0/src/std_measure.rs, a file whose header says "Generated by v1 compiler -- do not edit." The generator places a new coproduct in the .dag's own declaration order and sorts the `use` lines; I inserted PositiveSlotCount ahead of PositiveCelsiusDelta in both places. Same types, same functions, different bytes -- and the drift gate compares bytes, which is the entire point of it. Running claim_executor --required-regen --source-root dag --source-root src/v2 names std_measure.rs among the divergent surfaces, and its candidate differs from what I wrote by exactly those two reorderings and nothing else. The candidate is installed verbatim. I HAVE NOW MADE THIS MISTAKE TWICE IN ONE DAY. 2017773 regenerated this same mirror for TokensPerSecond and its message recorded why local verification cannot catch it -- every witness resolves .dag sources through an already-built binary, so the .dag change is live and the seed the next build needs is not. Adding a type to std/measure.dag has a mandatory second half, and knowing that is not the same as doing it. The actuator is cheap; running it is the check.
…1-fleet-serving-config # Conflicts: # src/v1/stage0/src/std_measure.rs
…ribing its figures Review 58311 on #9960. The note carried two byte figures and two context_length readings copied from a run, which DESIGN §6 forbids for the reason the neighbouring serving_desired.dag note already states: a number copied into prose is unreachable from the producer that owns it and rots without either end being touched. The qualitative conclusion is what the authority needs -- same context_length at one slot and at two, a larger buffer at two -- and the readback producer is named so the comparison is re-derived rather than read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
|
Fixed in Review 58311 is right and the fix is the one it names. The note in That is the same citation style the neighbouring — sent from eager-pike-541 |
…le this plan was in review The controlling reviewer found that main had falsified this plan's active DCH-0 claims. Verified here rather than relayed: #9960 (7810e68, an ancestor of current main) declares OllamaNumParallel in extdeps.ollama.server_env, binds ollama_num_parallel_env_assignment in spark.serving_unit_render so the rendered unit emits four axes rather than three, carries a desired slot count of 4, and corrects the cost model in the direction this plan predicted -- the count does not divide the window, and the per-slot price is recorded as a property of the realization under a declared 4b drop naming that subject. So the axis item and the cost-model item leave DCH-0. They are dispositioned in a new 2.1 "changes since the baseline" section rather than edited into the fixed main@de2f5f baseline, which stays immutable so later main movement does not rewrite a historical measurement. What #9960 does NOT resolve is kept open and sharpened. It made the state representable, declared and renderable; it read neither host back. Desired-versus- observed stays a live blocker, and the operator's two hand-edits make it sharper rather than weaker: a hand-set host value agreeing with a declared corpus value by coincidence is precisely what a silent carrier cannot distinguish from convergence, so the readback is the entire receipt. DCH-0 is renamed and its question narrowed to what actually remains: enrolment, and the live rendered-unit reconciliation. One defect found while verifying, reported and deliberately NOT repaired here because it is outside this lane: gunbc.spark.serving_desired still carries an earlier annotation asserting in the present tense that concurrency is not modeled in desired state, that the slot count divides the context window, and that the concurrency row is deliberately not smuggled in. All three are false, and they contradict that same module's own value and annotation about 140 lines below. A stale annotation is data the substrate cannot check, so nothing reds -- a 3 meaning fork inside a single authority. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6
…9977) * DCH-0 scope: the dedicated coding harness, and the four things that block it Scope doc only — no production rows, no types, no gate changes. Operator direction 2026-09-01: serve a large open model on the Sparks and drive a minimal coding harness against it, retiring the Claude/Codex/Cursor provider runtimes, written from the ground up in .dag and Rust. gunbc does not know about ctrl: no import, no dependency, no citation of a ctrl artifact as authority. The terminal is RLM's existing 14-step procedure with our harness selected as the provider realization. This lane authors no new acceptance procedure — RLM's is provider-agnostic in every step but one ("one provider process"), and reusing it is what keeps the harness from being graded on its own homework. Four blockers, each measured at the stated baseline rather than assumed: - The Sparks are NOT enrolled. fleet_intent_network.endpoints lists srv1-srv4 only, and srv5/srv6 are in neither it nor fleet_intent's ComputeHost list. In that module membership IS enrollment, so this is the reason parity is managed by hand. - "The models resident" is not representable. serving_desired governs one member; the materializer carries a single manifest, digest and blob closure, so no carrier could hold the seven the hosts actually serve. - Concurrency is unmodeled. serving_unit_render emits OLLAMA_HOST, OLLAMA_MODELS and OLLAMA_CONTEXT_LENGTH; OLLAMA_NUM_PARALLEL is absent, and unset means Ollama serializes. N concurrent harness sessions against one host is a queue. - There is no generic inference interface. extdeps/llm/llm.dag is 11 lines and llm_contracts.dag is 14; tool-calling is modeled only inside openai.dag and cursor_stream.dag, and anthropic.dag fuses shape with vendor. Probed live 2026-09-01, which is why the interface question has an answer: Ollama 0.32.9 on 192.168.1.225 answers POST /v1/messages with an Anthropic-shaped response carrying thinking and tool_use blocks and stop_reason "tool_use". No translating proxy is required. Ollama implementing the Anthropic wire shape is the shared-standard case DESIGN's external-upstream decomposition separates, so DCH-1 hoists the shape off the vendor rather than forking a third copy. Also recorded, because it is the class this subject keeps producing: the rendered unit is the only member of the must-move-together set with no failure signal. Ref, manifest, digest and closure each raise a checksum event when they disagree; a stale unit is silent, which is how both hosts sat with a Description naming gpt-oss:20b and no context-length line at all, unreported. DCH-0 (enrolment, residency, the parallel axis) is owned by eager-pike-541, to be confirmed with them before it starts. RLM closes first. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * DCH-0 restructure: split residency into its own gate, and the endpoints census says list membership is not the load-bearing act Three changes, all from eager-pike-541's reply plus one census I owed them. SPLIT. Residency modeling becomes DCH-0b. Their reason is the right one and it is not about size: enrolment and the parallel axis write values into carriers that already exist and have shapes, while the resident set has NO carrier at all -- what a resident-model fact even is, desired or observed, per host or per fleet, its identity when the same weights appear under two refs, remains undecided. Folding design into a gate of row-writes makes the row-writes wait on it. THE CENSUS THEY FLAGGED AS NOT DONE IS DONE, and it changes DCH-0's content. Of 57 non-test modules importing fleet_intent_network, the number reading the endpoints list or fleet_intent_network_topology() is ZERO -- production consumers take individual endpoint rows by name (bmc_virtual_media, host_identity_access). The topology function has exactly one consumer in the tree: the witness asserting list_length(endpoints) == 11. So list membership is not the load-bearing act; authoring the srv5/srv6 rows that named consumers can reach is. And enrolment turns that witness red at 11 -> 13, where the right response is not to bump the literal -- a count copied from the tree it measures is the change detector DESIGN section 5 names, and completeness is an identity join. Repairing that oracle is now part of the gate, and bumping it is on the must-not list. THE COST MODEL IN THE CORPUS IS FALSE and DCH-0 now corrects it while adding the axis. serving_desired says context length and slot count "trade against each other inside one memory budget, because the slot count divides the context window." Measured by eager-pike-541 on idle spark-3bd5, same model, only slots changed: 1 slot -> context_length 1048576 at size_vram 88,865,253,620; 2 slots -> the same 1048576 at 90,543,761,652. The window is not divided. Slot count multiplies KV. The sentence is deleted rather than softened, and the measured per-slot cost goes behind the desired value -- with the caveat that 1.68 GB/slot is a DeepSeek MLA number and a fleet-wide slot count from one realization is the same overreach as a fleet-wide context ceiling from one realization. Also recorded: enrolment is inert because it has no executor -- no ctrl-fleet-converge timer or service on either Spark in either scope, and fleet_converge_timer says of itself that it is no longer a renderer since #8283 killed its installer chain. The risk is therefore not prematurity but a row that reads as an outcome, which is why the receipt stays an observed-vs-desired rendered-unit comparison. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * Enrolment is two acts and only one is inert: ComputeHost has three consumers and one of them is a fail-closed wall The endpoints census generalized further than it should have. eager-pike-541 censused the other list and it is the opposite shape; I verified the load-bearing consumer in the tree before folding it in. fleet_intent_known_hosts has three production consumers: - gunbc.generated_artifact maps it to RunnerHostSudoersArtifact, so enrolment mints two new generated artifacts and the rows cannot land without a same-commit regeneration or the drift gate reds. - runner_host_deploy admit_runner_host returns RunnerHostUnenrolled for absent hosts. Its annotation calls itself THE ENROLLMENT WALL and is explicit that it is CONSTRUCTION and not a check -- no function in the module takes a bare RunnerHostDeploy and yields a command, so running an installer on an unenrolled host is unrepresentable rather than discouraged. It prices the srv4 host-convergence OOM and names the enrolled row as where runner_deployment_plan's conservation wall gets host RAM to check slot caps against. Enrolling srv5/srv6 therefore REMOVES a fail-closed refusal that currently protects them. - fleet_converge_apply finds hosts by identity in that list, so enrolment is what makes a Spark reachable by the apply path this lane measured as PartiallyApplied and refused. The inertness argument had two independent legs, no executor and no consumer. ComputeHost knocks out the second; only the executor leg survives, and it is the leg that changes the moment anyone installs the timer. So DCH-0 lands the endpoints half only. The ComputeHost half becomes DCH-0c and is a decision rather than a row: host facts the conservation wall consumes, the two sudoers artifacts regenerated in the same commit, and someone saying out loud that the wall is coming down for these machines on purpose. The row/list split holds on both sides and inverts between them -- for endpoints the row is load-bearing and the list inert; for ComputeHost the list is load-bearing and the row quiet. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * Narrow the ctrl boundary to the lane: the global form was false of the repository Review found a blocking defect in the opening direction. It read "gunbc does not know about ctrl -- no import, no dependency, no citation of a ctrl artifact as authority", which is globally false: extdeps.ctrl.gunbc_pin already declares an ExternalAuthority whose URI points into gunb-ai/ctrl, and its functions compare the host pin against the ctrl pin. The operator's intent was about DCH and the harness, and the document already stated it correctly at that grain further down, under "What this lane must not do". So the opening asserted a stronger claim than the one being made, and in an authority-bearing scope document where independence from ctrl is a central boundary, a false global subject is not harmless prose -- a later reader would take it as a fact about the repository and find a counterexample in one grep. Narrowed to the lane, with the counterexample named in place so the narrowing cannot be re-widened by someone who does not know why it is narrow, and with an explicit note that extdeps.ctrl.gunbc_pin is out of scope here and not to be touched. No other plan content moves; this is the only authored-content change from fddf116. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * DCH: answer the relayed harness transcript as an adversarial review, and record the terminal as necessary-not-sufficient The operator relayed a transcript of the ctrl mini-agent lane debugging its own harness against the same Sparks DCH intends to drive, and directed that the plan answer it before the lane starts. Admitted on the terms the plan already sets for that lane's output: second-hand, unreproduced here, each row a hypothesis about our design with a predicted failure rather than a fact on loan. What it buys is not findings but a cheap enumeration of where a harness of this shape breaks. Nine classes, each routed to the gate that owns it. They collapse to one property: every failure was operator-visible AS SOMETHING OTHER THAN ITSELF -- a harness throw as a clean exit, a dead session as a working one, a control-plane restart as a Spark fault, an inflated rate as fast hardware. So the obligation is not "handle these bugs" but produce a disposition that cannot be mistaken for a different one. Three consequences worth naming outside the table: - The plan's DCH-2 claim that a self-reporting harness retires the observation layer now has its measurement. Their harness DID report its status correctly throughout, and the operator still saw a frozen session, because the report's transport failed independently of the report. Self-reporting is necessary and is not the simplification alone. - DCH-0 puts the serving unit under fleet convergence, and a converge restarts the serving process -- so this lane imports that transcript's worst open defect BY CONSTRUCTION, into the layer it chose on purpose. Sent for adjudication with a position rather than decided here. - The terminal is amended to necessary-without-being-sufficient. The canary will not fill a context window, meet a tool timeout, or sit through a converge, so a green on it is a green on the easy path. No substitute terminal still counts. Five questions are open with the controlling reviewer and the section is not settled until they return; three of them can change a gate. Two new hypotheses join the existing list on the same terms -- streaming-specific early termination, and a turn ending at end_turn with announced work unperformed. Neither is established: the second has one observation under survivorship, the first has none, because the step isolating streaming as the single variable was never run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * DCH: reconcile against #9960, which resolved the concurrency axis while this plan was in review The controlling reviewer found that main had falsified this plan's active DCH-0 claims. Verified here rather than relayed: #9960 (7810e68, an ancestor of current main) declares OllamaNumParallel in extdeps.ollama.server_env, binds ollama_num_parallel_env_assignment in spark.serving_unit_render so the rendered unit emits four axes rather than three, carries a desired slot count of 4, and corrects the cost model in the direction this plan predicted -- the count does not divide the window, and the per-slot price is recorded as a property of the realization under a declared 4b drop naming that subject. So the axis item and the cost-model item leave DCH-0. They are dispositioned in a new 2.1 "changes since the baseline" section rather than edited into the fixed main@de2f5f baseline, which stays immutable so later main movement does not rewrite a historical measurement. What #9960 does NOT resolve is kept open and sharpened. It made the state representable, declared and renderable; it read neither host back. Desired-versus- observed stays a live blocker, and the operator's two hand-edits make it sharper rather than weaker: a hand-set host value agreeing with a declared corpus value by coincidence is precisely what a silent carrier cannot distinguish from convergence, so the readback is the entire receipt. DCH-0 is renamed and its question narrowed to what actually remains: enrolment, and the live rendered-unit reconciliation. One defect found while verifying, reported and deliberately NOT repaired here because it is outside this lane: gunbc.spark.serving_desired still carries an earlier annotation asserting in the present tense that concurrency is not modeled in desired state, that the slot count divides the context window, and that the concurrency row is deliberately not smuggled in. All three are false, and they contradict that same module's own value and annotation about 140 lines below. A stale annotation is data the substrate cannot check, so nothing reds -- a 3 meaning fork inside a single authority. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * DCH: carry the open dispositions into the gate chain, and specify the tool surface instead of forward-referencing it Review 58352 (codex/gpt-5.6-sol) requested changes on two findings. Both are correct and both are repaired here rather than argued. FIRST: the review section identified five gate-changing questions and left them beside the gate chain rather than inside it, so DCH-0 could begin and complete while knowingly preserving the restart-mid-turn fail-open that Q3 names. Recording a class is not handling it, and 5 exists to answer these in advance. The gate chain now carries an admission rule -- a gate does not complete while a 5.2 question owning one of its dispositions is open -- with the bindings tabled and repeated in each gate. DCH-0 states its own ending explicitly: a drain or lease so a converge cannot land against an in-flight turn, or a 4b(3) row whose trigger names the drain capability. Never silence, and never a note that a restart is unlikely. The rule bars a gate's RECEIPT, not work inside it. SECOND: the section said the tool-surface quirks were "expanded there rather than here" while DCH-2 still listed four tool names. That is a promise naming a route that does not exist -- the same defect this plan refuses in others, committed by me in the act of writing the section that refuses it. DCH-2 now specifies the surface as six owed dispositions under one rule: a tool's contract is carried in its declared shape, and any limit it enforces is reported in its result rather than inferred from a truncated or missing one. Working directory, duration, input, truncation, edit matching, context exhaustion. Two of the relayed harness's choices are adopted deliberately because they are already the right shape -- a cap that announces its truncation, and an edit that refuses on zero or multiple matches rather than patching the first. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * DCH: carry the five returned dispositions into the gates, and add DCH-0r All five questions returned 2026-09-02. Recorded as rulings rather than positions; where a proposal was accepted with refinement, the refinement binds. The chain changes. Q3 rejected a 4b drop as the normal shape -- a drop reports a lower guarantee, it does not turn an unsafe arm green -- so the drain becomes its own gate rather than debt, and reordering removes the need for debt entirely: DCH-0 endpoint rows -> DCH-0r quiescent maintenance -> DCH-0c ComputeHost enrolment -> DCH-2 turn activation DCH-0r's mechanical blocker is verified in this tree rather than relayed, and verifying it found two refinements sharper than the finding as received. spark.serving_realization realizes EnableSystemUnit as enable, start AND restart in one effect, so restart has no independently matchable identity and no fence can guard it. First: that restart is UNCONDITIONAL, emitted whenever the effect is realized rather than when the definition changed, so converging the unit is destructive to in-flight work even when it changes nothing. Second, and the Sparks are on this path: the USER-unit closure already separates EnableUserUnit from StartUserUnit and carries NO restart effect at all -- so on the hosts this lane targets, a changed unit definition has no modeled route to take effect. That is a candidate cause for the observed-versus-desired drift DCH-0 must reconcile, and it is written as a candidate because nothing here has read a host back. Also carried: the admission fence and the drain are ONE protocol, because a bare active==0 read admits a turn immediately after it; and the guarantee grain is stated rather than assumed, since a census of DCH leases proves quiescence only for DCH-managed clients. DCH-2 gains an exit bar of four qualification groups and no longer exits on "four tools and one terminal turn". DCH-3 requires BOTH the qualification receipts and the unchanged 14-step terminal, neither substituting for the other. DCH-4 waits on that combined admission, because retiring a working runtime on an easy-path green is how a replacement erases a correctness distinction instead of completing. Q1 refines my boundary: whose failure is evidence, but the smallest closed causal scope containing the uncertainty decides the stop, and a supervisor alive in a typed stopped state is not an absorbing fallback. Q2 is accepted at typed- obligation grain and rejected at prose or character grain, with awaiting- verification deliberately not acceptance since the model's own EndTurn cannot grade the model's work. Q4 holds n=1 under survivorship to a single existential proposition that prior art does not establish for our harness at all. Q5 replaces both my options with two axes plus journal-before-effect ordering: a supervisor may move a record from running to interrupted and never to accepted. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * DCH: adopt the operator's only-and-default direction, and separate default from retirement Operator direction 2026-09-02: make the .dag harness the only and default harness, ignoring Claude and Codex entirely. Adopted, with the two halves kept apart because they have different preconditions. DEFAULT IS A SELECTION; RETIREMENT IS A DELETION. "Ignore the other providers" is discharged by making ours the default selection at the seam and stopping investment in the others. It does not require deleting them, and DCH-4's deletion still waits on DCH-3's receipt. Deleting them to satisfy "only" before that receipt would leave the repository with zero working harnesses -- the failure the replacement-migration doctrine calls erasing a correctness distinction rather than completing the replacement. The other variants stay frozen in the doctrine's sense until DCH-4. "ONLY" RAISES THE BAR RATHER THAN LOWERING IT. With three providers a weak DCH-3 is tolerable because a fallback exists; with one there is none, so every DCH-2 qualification class becomes load-bearing on the day it becomes default. The relayed transcript in 5 is exactly a record of what a sole harness with no fallback feels like when it breaks: three sessions dead, an exit status of zero, and an operator who cannot tell working from dead. The combined admission is what makes "only" survivable, so it is not tradeable for focus. Also added to the must-not-do list: do not delete or break the existing runtimes to satisfy "only" before DCH-3's receipt, or the count of working harnesses passes through zero. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * DCH: withdraw the RLM-first start barrier, and correct DCH-0r's identity model to three identities The reviewer superseded my freeze at ef8cf03 for a reason I should have caught myself: the plan still contained the OLD sequencing in two places -- the opening "RLM closes first; this lane starts after it" and the section 4 prohibition on starting before RLM's terminal receipt -- while DCH-3 had already been rewritten to the operator's 2026-09-02 direction. The document contradicted itself, and no freeze can make contradictory plan text admissible. Concurrent DCH and RLM is authorized. DCH-3 is a JOIN, not a start barrier: DCH-0 through DCH-2 build and qualify while RLM is blocked, and the default cutover waits for both sides. "Only" raises that bar rather than licensing an early one. DCH-0r's protocol carried two conflations and the correction is now the load-bearing part of the gate. THREE identities, not one: a stable service locus, a service-invocation identity that changes on every activation, and a running-definition identity stable for equal specs. The /proc environ stamp answers the THIRD -- it is invocation-bound evidence but not an invocation identity, because two successive restarts from one definition carry the same stamp and so it cannot prove the old process was replaced. And the fence is keyed by the LOCUS, not the incarnation. An incarnation-scoped fence stops blocking admissions at exactly the moment an unvalidated new incarnation appears mid-drain. The pre-restart invocation is a CAS condition and an evidence anchor, never the scope of the refusal. The unstamped reading is narrowed. "No stamp" establishes that the invocation cannot be shown to have started from the current desired definition; it does NOT establish that the process predates it, since a foreign or manually started process could have been launched later and would carry no stamp either. It still requires reactivation; chronology is just not what the evidence supports. Also added to the must-not-do list: do not land a change to an authority already inside another lane's frozen approved delta without serializing it. A clean textual composition still yields a blob nobody approved. dag/gunbc/fleet/fleet_converge_plan.dag is the live instance, shared with #9832. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * DCH: model the behaviour-bearing serving package, add DCH-0p, and retire the degenerate-stop hypothesis to a mechanical cause Ruling 2026-09-02. The plan's model identity was weights-shaped and that is insufficient: same weight blob, same template, same runtime, same host, DIFFERENT EFFECTIVE STOP SET is materially different observable behaviour. So DCH-0b's subject becomes a ServingModelPackageIdentity carrying template, parser, parameters, the effective stop-sequence set, runtime release and decode realization, with a tag as an allocated handle and never the semantic identity. A package rebuilt from the same weights without a stop string is a DIFFERENT package; treating it as identical makes the repair invisible to convergence. New gate DCH-0p, after DCH-0b and DCH-0r and before DCH-2 qualification. It reads what the endpoint ACTUALLY serves rather than what a Modelfile or a tag says it serves, refuses with PackageConfigurationIncomplete rather than inferring a behaviour-bearing field, and admits a stop sequence only when it binds uniquely to a declared protocol boundary in that exact package's contract. A copied allowlist is not a role, and sentinel-looking values inherited from another lane are not our evidence. THE HYPOTHESIS THIS RETIRES IS THE REASON THE LANE EXISTS. 5 recorded, from the relayed transcript, that an assistant turn may stop at end_turn having announced work it did not perform -- one observation, survivorship-filtered, which I was careful to say established almost nothing. It now has a mechanical cause: a serving package carrying PARAMETER stop ### terminates generation on ordinary Markdown H3 output while returning a well-formed stream and stop_reason end_turn. Reproduced by a tiny prompt, removed by rebuilding the package. So the heading-then-stop symptom had a determinate external cause the whole time, and the hypothesis about model behaviour was a transport symptom read as a semantic one. It leaves the hypothesis list and becomes DCH-0p's subject. The streaming hypothesis is untouched and still has NO observation. DCH-2 gains two things. Its terminal semantics may never equate end_turn with natural completion, because at least two upstream causes produce that same label and the client cannot prove the native cause from a surface that erased it -- so DCH-0p and DCH-2 are two independent walls rather than one. And throughput becomes a construction rather than a threshold: a cache lineage joined to accounting that separates newly evaluated from reused input tokens, since a rate over total input divides mismatched subjects and yields a valid arithmetic operation that is not a hardware measurement. "No four-digit rate is reachable" is a view, which #9946 forbids as a correctness wall. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 --------- Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Three fleet-serving facts that were misnamed, missing, or wrong. Split out of #9897, which is a model-authority PR and was carrying this only because both edits happened in the same hour.
Misnamed
extdeps.ollama.server_envcalledOLLAMA_CONTEXT_LENGTH"the context window the SERVER allocates" and "the CONFIGURED limit". Upstream 0.32.9 resolves options in layers — environment, then model, then request — and an explicitnum_ctxat either later layer replaces it. It never limited anything: it is the value used when no override applies. One name carrying two materially different contracts is a meaning fork, and a deployment reading it as a ceiling would believe it had bounded something any caller can raise. Renamed through the carriers todefault_context.Missing
OLLAMA_NUM_PARALLELwas not in the closed variable set and not rendered, so the hosts inherited whatever the runtime chose — one slot, which serializes. Measured 2026-09-01: a ~61,000-token prefill held one host for roughly two hundred seconds at a healthy 283–331 tok/s while every other caller queued, and the queue was read as a latency regression in the model rather than as contention. The axis is now modeled and the desired value is 4.Wrong
The P1b note said the two axes trade "because the slot count divides the context window." Measured on one host, one model, changing only that variable:
context_lengthsize_vramEach slot receives its own full window. The count multiplies memory; it does not divide context. Deleted rather than softened.
The desired context becomes 1,048,576, and its evidence sentence is replaced
The old one cited an
/api/psreading as proof of a configured value. But/api/psreports what a loaded runner holds, and both hosts carried noOLLAMA_CONTEXT_LENGTHline at all — so the number being confirmed was the runtime's own automatic default. A declaration cited its own absence as confirmation. Two prior values, 8192 and 131072, were both never enacted.The rung drop is filed
One row for both values, because they are one defect in two spellings: measured on one node, one artifact, one runtime, one concurrency condition — then applied to two hosts and every realization the cell loads.
llama4:scoutin particular declares a window far beyond this ceiling with no cache compression.The trigger names the capability, not an artifact: an observation table alone does not discharge it, because while a global default exists an unmeasured realization still inherits these numbers automatically. An unobserved realization must refuse, not inherit.
Standing disclosure
Both values are already live on both hosts, set by hand under operator instruction while they were blocked on testing. The modeled route —
fleet_converge_timer → fleet_converge_apply → host_effect_realize— exists and carries spark serving rows, but was not usable at the time: thegunbc convergeCLI verb is unwired by design, and a peer lane's wet plan run on that spine returnedPartiallyAppliedwith apply refused. This PR reconciles the corpus to the hosts; until it lands, canonical main selects 131,072 and the hosts are intentional drift.