Repository navigation
File selection_view_read_as_population: a threshold over measured values is a view, and it arrives as the fix for a real objection - #9946
Conversation
…ues is a view, and it arrives as the fix for a real objection A decidable predicate is written over MEASURED values and its output is then treated as the population of the class under study. Membership in that set is a property of the measurement, not of the subject: it moves when the instrument or the load moves while the subject is unchanged. What separates it from its two neighbours is the route it arrives by. instrument_output_read_as_subject_content is a report read past its own promise and censored_estimator_drops_its_own_tail is a statistic over survivors -- both are errors in the reading. This one is produced BY diligence: the predicate is decidable, honestly derived, and usually the correct answer to a legitimate refusal, which is exactly why nobody asks whether it answers the question that was posed. Being careful is what makes it invisible. Both receipts belong to the same author three hours apart, which is why it is filed rather than blamed: naming the truncated-top-25 defect in a peer's median, then building a "measured cpu >= N" admission rule -- the correct repair for a censored-parameter refusal -- and writing its output into a declared drop as the population. Carries the monotonicity corollary as a separate seam: the evidence floor behind such a threshold may be monotone while the membership set is not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XpYuLQ4AwfSJmcFk26aPap
…ting it by identity does not resolve (review 58237) The row contrasted itself with censored_estimator_drops_its_own_tail, which lives in an unmerged PR -- so on this branch, and on main, the citation resolves to nothing. A canonical row asserting a boundary against a nonexistent authority is exactly the unresolvable-citation defect the namespace discipline exists to prevent, and it is worse than no citation because a reader takes the name as established. The boundary is now drawn by SHAPE -- a statistic computed over the observations that survived a threshold on the very variable being estimated -- with the sibling named as filed in gunbc#9932 and explicitly not cited by identity while it is unlanded. instrument_output_read_as_subject_content, which does resolve on main, is still cited by name. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XpYuLQ4AwfSJmcFk26aPap
|
Addressed review 58237 (REQUEST_CHANGES) — the finding is correct and the citation is gone. The defect: this row drew its load-bearing distinction against What it was: I branched this PR off main deliberately, to keep a second class out of #9932 (which had been re-authored twice under review, where a rider gets skimmed rather than read). That was the right call for the PR boundary and it is exactly what created the dangling reference — the sibling row was in my head and in another branch, not in the tree this row will land in. The fix: the boundary is now drawn by shape rather than by identity — "a statistic computed over the observations that survived a threshold on the very variable being estimated" — with the sibling named as filed in gunbc#9932 and explicitly marked as not cited by identity while it is unlanded. What I did not do, and why: I did not make this PR depend on #9932 landing first. The two are independent classes and stacking them would put this row behind a PR that has already been re-authored twice; if #9932 lands first, the descriptive sentence can be tightened to a symbol citation in a one-line follow-up, and if it does not land, this row is still correct as written. A citation that resolves only in a hypothetical future tree is the same defect deferred. — sent from deep-badger-41 |
…1-selection-view # Conflicts: # DESIGN.md # docs/design-ledgers.md
…1-selection-view # Conflicts: # DESIGN.md # dag/gunbc/recurring_failure_mode.dag # docs/design-ledgers.md
… 58299) The specimen said a peer computed statistics over "the 25 rows a floor log printed" -- a transcribed number with no producer, which DESIGN section 6 refuses: a measurement is cited by naming the thing that re-derives it, never by copying its output, because a transcribed figure is unreachable from the instrument that owns it. Both producers are now named and no count is asserted: the truncated view is a required-floor run's [over-cost] listing, which ranks by cost and truncates to a fixed head and whose own closing line says so, and the population is required_floor_claim_cost.tsv as uploaded by the same run. Either is re-derivable from any floor run. That the row about a truncated view carried an untraceable number FROM that view is not lost on me; naming the instrument is also what lets a reader check the specimen rather than believe it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XpYuLQ4AwfSJmcFk26aPap
|
Addressed review 58299 (REQUEST_CHANGES) — the finding is right and the irony is not lost on me. The specimen said a peer computed statistics over "the 25 rows a floor log printed." That is a transcribed number with no producer, which §6 refuses: a measurement is cited by naming the thing that re-derives it, never by copying its output, because a transcribed figure is unreachable from the instrument that owns it and rots without either end being touched. Both producers are now named and no count is asserted. The truncated view is a required-floor run's Worth stating plainly: this is a row about a truncated view being mistaken for a population, and it carried an untraceable number taken from that very view. Naming the instrument is not just §6 compliance here — it is what lets a reader check the specimen instead of trusting it, which is the row's own point turned on the row. — sent from deep-badger-41 |
Roster conflict resolved additively: main's censored_estimator_drops_its_own_tail (landed with #9932) and this branch's selection_view_read_as_population both kept. Both doc projections reset to main's side and regenerated narrowly per path with main_wet_one; the resulting diff against main is two insertions and zero deletions. The row's boundary clause is updated for the fact the merge creates: it carried censored_estimator_drops_its_own_tail DESCRIPTIVELY while that row was unlanded, and now cites it by identity, because in this tree the symbol resolves.
Main moved to 99ace7a (#9946, selection_view_read_as_population), which touched DESIGN.md and docs/design-ledgers.md — the same two generated projections this branch touches. GitHub reported mergeable=CLEAN, which is not evidence: it does not run this repository's generated-artifact merge driver, so it reports a clean TEXT merge on projections whose bytes would then project neither side's authorities. `git merge-tree --write-tree` is the authority, and it refused with GeneratedArtifactConcurrentDivergence on both paths (gunbc#9969). Neither projection is hand-resolved. The driver left both UNMERGED with the ours side verbatim and no conflict markers, and both were REGENERATED from the merged authorities via `dag/gunbc/instruments/generated_artifact_gate.dag main_wet`. Both ledger rows survive the merge: this branch's `compensating_errors_cancel_in_the_aggregate` and main's `selection_view_read_as_population`. Also adds the boundary sentence the class needed: the row now states explicitly that NO SWEEP FOR SIBLINGS WAS PERFORMED, so the absence of a census reads as a declared boundary rather than as coverage. The receipt establishes the class at one subject and says nothing about the population — reading the row as a census of aggregate-cardinality checks would be the same substitution it names. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P7mphvNU1JoCbrowqDM5Zg
…ire the degenerate-stop hypothesis to a mechanical cause Ruling 2026-09-02. The plan's model identity was weights-shaped and that is insufficient: same weight blob, same template, same runtime, same host, DIFFERENT EFFECTIVE STOP SET is materially different observable behaviour. So DCH-0b's subject becomes a ServingModelPackageIdentity carrying template, parser, parameters, the effective stop-sequence set, runtime release and decode realization, with a tag as an allocated handle and never the semantic identity. A package rebuilt from the same weights without a stop string is a DIFFERENT package; treating it as identical makes the repair invisible to convergence. New gate DCH-0p, after DCH-0b and DCH-0r and before DCH-2 qualification. It reads what the endpoint ACTUALLY serves rather than what a Modelfile or a tag says it serves, refuses with PackageConfigurationIncomplete rather than inferring a behaviour-bearing field, and admits a stop sequence only when it binds uniquely to a declared protocol boundary in that exact package's contract. A copied allowlist is not a role, and sentinel-looking values inherited from another lane are not our evidence. THE HYPOTHESIS THIS RETIRES IS THE REASON THE LANE EXISTS. 5 recorded, from the relayed transcript, that an assistant turn may stop at end_turn having announced work it did not perform -- one observation, survivorship-filtered, which I was careful to say established almost nothing. It now has a mechanical cause: a serving package carrying PARAMETER stop ### terminates generation on ordinary Markdown H3 output while returning a well-formed stream and stop_reason end_turn. Reproduced by a tiny prompt, removed by rebuilding the package. So the heading-then-stop symptom had a determinate external cause the whole time, and the hypothesis about model behaviour was a transport symptom read as a semantic one. It leaves the hypothesis list and becomes DCH-0p's subject. The streaming hypothesis is untouched and still has NO observation. DCH-2 gains two things. Its terminal semantics may never equate end_turn with natural completion, because at least two upstream causes produce that same label and the client cannot prove the native cause from a surface that erased it -- so DCH-0p and DCH-2 are two independent walls rather than one. And throughput becomes a construction rather than a threshold: a cache lineage joined to accounting that separates newly evaluated from reused input tokens, since a rate over total input divides mismatched subjects and yields a valid arithmetic operation that is not a hardware measurement. "No four-digit rate is reachable" is a view, which #9946 forbids as a correctness wall. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6
…9977) * DCH-0 scope: the dedicated coding harness, and the four things that block it Scope doc only — no production rows, no types, no gate changes. Operator direction 2026-09-01: serve a large open model on the Sparks and drive a minimal coding harness against it, retiring the Claude/Codex/Cursor provider runtimes, written from the ground up in .dag and Rust. gunbc does not know about ctrl: no import, no dependency, no citation of a ctrl artifact as authority. The terminal is RLM's existing 14-step procedure with our harness selected as the provider realization. This lane authors no new acceptance procedure — RLM's is provider-agnostic in every step but one ("one provider process"), and reusing it is what keeps the harness from being graded on its own homework. Four blockers, each measured at the stated baseline rather than assumed: - The Sparks are NOT enrolled. fleet_intent_network.endpoints lists srv1-srv4 only, and srv5/srv6 are in neither it nor fleet_intent's ComputeHost list. In that module membership IS enrollment, so this is the reason parity is managed by hand. - "The models resident" is not representable. serving_desired governs one member; the materializer carries a single manifest, digest and blob closure, so no carrier could hold the seven the hosts actually serve. - Concurrency is unmodeled. serving_unit_render emits OLLAMA_HOST, OLLAMA_MODELS and OLLAMA_CONTEXT_LENGTH; OLLAMA_NUM_PARALLEL is absent, and unset means Ollama serializes. N concurrent harness sessions against one host is a queue. - There is no generic inference interface. extdeps/llm/llm.dag is 11 lines and llm_contracts.dag is 14; tool-calling is modeled only inside openai.dag and cursor_stream.dag, and anthropic.dag fuses shape with vendor. Probed live 2026-09-01, which is why the interface question has an answer: Ollama 0.32.9 on 192.168.1.225 answers POST /v1/messages with an Anthropic-shaped response carrying thinking and tool_use blocks and stop_reason "tool_use". No translating proxy is required. Ollama implementing the Anthropic wire shape is the shared-standard case DESIGN's external-upstream decomposition separates, so DCH-1 hoists the shape off the vendor rather than forking a third copy. Also recorded, because it is the class this subject keeps producing: the rendered unit is the only member of the must-move-together set with no failure signal. Ref, manifest, digest and closure each raise a checksum event when they disagree; a stale unit is silent, which is how both hosts sat with a Description naming gpt-oss:20b and no context-length line at all, unreported. DCH-0 (enrolment, residency, the parallel axis) is owned by eager-pike-541, to be confirmed with them before it starts. RLM closes first. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * DCH-0 restructure: split residency into its own gate, and the endpoints census says list membership is not the load-bearing act Three changes, all from eager-pike-541's reply plus one census I owed them. SPLIT. Residency modeling becomes DCH-0b. Their reason is the right one and it is not about size: enrolment and the parallel axis write values into carriers that already exist and have shapes, while the resident set has NO carrier at all -- what a resident-model fact even is, desired or observed, per host or per fleet, its identity when the same weights appear under two refs, remains undecided. Folding design into a gate of row-writes makes the row-writes wait on it. THE CENSUS THEY FLAGGED AS NOT DONE IS DONE, and it changes DCH-0's content. Of 57 non-test modules importing fleet_intent_network, the number reading the endpoints list or fleet_intent_network_topology() is ZERO -- production consumers take individual endpoint rows by name (bmc_virtual_media, host_identity_access). The topology function has exactly one consumer in the tree: the witness asserting list_length(endpoints) == 11. So list membership is not the load-bearing act; authoring the srv5/srv6 rows that named consumers can reach is. And enrolment turns that witness red at 11 -> 13, where the right response is not to bump the literal -- a count copied from the tree it measures is the change detector DESIGN section 5 names, and completeness is an identity join. Repairing that oracle is now part of the gate, and bumping it is on the must-not list. THE COST MODEL IN THE CORPUS IS FALSE and DCH-0 now corrects it while adding the axis. serving_desired says context length and slot count "trade against each other inside one memory budget, because the slot count divides the context window." Measured by eager-pike-541 on idle spark-3bd5, same model, only slots changed: 1 slot -> context_length 1048576 at size_vram 88,865,253,620; 2 slots -> the same 1048576 at 90,543,761,652. The window is not divided. Slot count multiplies KV. The sentence is deleted rather than softened, and the measured per-slot cost goes behind the desired value -- with the caveat that 1.68 GB/slot is a DeepSeek MLA number and a fleet-wide slot count from one realization is the same overreach as a fleet-wide context ceiling from one realization. Also recorded: enrolment is inert because it has no executor -- no ctrl-fleet-converge timer or service on either Spark in either scope, and fleet_converge_timer says of itself that it is no longer a renderer since #8283 killed its installer chain. The risk is therefore not prematurity but a row that reads as an outcome, which is why the receipt stays an observed-vs-desired rendered-unit comparison. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * Enrolment is two acts and only one is inert: ComputeHost has three consumers and one of them is a fail-closed wall The endpoints census generalized further than it should have. eager-pike-541 censused the other list and it is the opposite shape; I verified the load-bearing consumer in the tree before folding it in. fleet_intent_known_hosts has three production consumers: - gunbc.generated_artifact maps it to RunnerHostSudoersArtifact, so enrolment mints two new generated artifacts and the rows cannot land without a same-commit regeneration or the drift gate reds. - runner_host_deploy admit_runner_host returns RunnerHostUnenrolled for absent hosts. Its annotation calls itself THE ENROLLMENT WALL and is explicit that it is CONSTRUCTION and not a check -- no function in the module takes a bare RunnerHostDeploy and yields a command, so running an installer on an unenrolled host is unrepresentable rather than discouraged. It prices the srv4 host-convergence OOM and names the enrolled row as where runner_deployment_plan's conservation wall gets host RAM to check slot caps against. Enrolling srv5/srv6 therefore REMOVES a fail-closed refusal that currently protects them. - fleet_converge_apply finds hosts by identity in that list, so enrolment is what makes a Spark reachable by the apply path this lane measured as PartiallyApplied and refused. The inertness argument had two independent legs, no executor and no consumer. ComputeHost knocks out the second; only the executor leg survives, and it is the leg that changes the moment anyone installs the timer. So DCH-0 lands the endpoints half only. The ComputeHost half becomes DCH-0c and is a decision rather than a row: host facts the conservation wall consumes, the two sudoers artifacts regenerated in the same commit, and someone saying out loud that the wall is coming down for these machines on purpose. The row/list split holds on both sides and inverts between them -- for endpoints the row is load-bearing and the list inert; for ComputeHost the list is load-bearing and the row quiet. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * Narrow the ctrl boundary to the lane: the global form was false of the repository Review found a blocking defect in the opening direction. It read "gunbc does not know about ctrl -- no import, no dependency, no citation of a ctrl artifact as authority", which is globally false: extdeps.ctrl.gunbc_pin already declares an ExternalAuthority whose URI points into gunb-ai/ctrl, and its functions compare the host pin against the ctrl pin. The operator's intent was about DCH and the harness, and the document already stated it correctly at that grain further down, under "What this lane must not do". So the opening asserted a stronger claim than the one being made, and in an authority-bearing scope document where independence from ctrl is a central boundary, a false global subject is not harmless prose -- a later reader would take it as a fact about the repository and find a counterexample in one grep. Narrowed to the lane, with the counterexample named in place so the narrowing cannot be re-widened by someone who does not know why it is narrow, and with an explicit note that extdeps.ctrl.gunbc_pin is out of scope here and not to be touched. No other plan content moves; this is the only authored-content change from fddf116. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * DCH: answer the relayed harness transcript as an adversarial review, and record the terminal as necessary-not-sufficient The operator relayed a transcript of the ctrl mini-agent lane debugging its own harness against the same Sparks DCH intends to drive, and directed that the plan answer it before the lane starts. Admitted on the terms the plan already sets for that lane's output: second-hand, unreproduced here, each row a hypothesis about our design with a predicted failure rather than a fact on loan. What it buys is not findings but a cheap enumeration of where a harness of this shape breaks. Nine classes, each routed to the gate that owns it. They collapse to one property: every failure was operator-visible AS SOMETHING OTHER THAN ITSELF -- a harness throw as a clean exit, a dead session as a working one, a control-plane restart as a Spark fault, an inflated rate as fast hardware. So the obligation is not "handle these bugs" but produce a disposition that cannot be mistaken for a different one. Three consequences worth naming outside the table: - The plan's DCH-2 claim that a self-reporting harness retires the observation layer now has its measurement. Their harness DID report its status correctly throughout, and the operator still saw a frozen session, because the report's transport failed independently of the report. Self-reporting is necessary and is not the simplification alone. - DCH-0 puts the serving unit under fleet convergence, and a converge restarts the serving process -- so this lane imports that transcript's worst open defect BY CONSTRUCTION, into the layer it chose on purpose. Sent for adjudication with a position rather than decided here. - The terminal is amended to necessary-without-being-sufficient. The canary will not fill a context window, meet a tool timeout, or sit through a converge, so a green on it is a green on the easy path. No substitute terminal still counts. Five questions are open with the controlling reviewer and the section is not settled until they return; three of them can change a gate. Two new hypotheses join the existing list on the same terms -- streaming-specific early termination, and a turn ending at end_turn with announced work unperformed. Neither is established: the second has one observation under survivorship, the first has none, because the step isolating streaming as the single variable was never run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * DCH: reconcile against #9960, which resolved the concurrency axis while this plan was in review The controlling reviewer found that main had falsified this plan's active DCH-0 claims. Verified here rather than relayed: #9960 (7810e68, an ancestor of current main) declares OllamaNumParallel in extdeps.ollama.server_env, binds ollama_num_parallel_env_assignment in spark.serving_unit_render so the rendered unit emits four axes rather than three, carries a desired slot count of 4, and corrects the cost model in the direction this plan predicted -- the count does not divide the window, and the per-slot price is recorded as a property of the realization under a declared 4b drop naming that subject. So the axis item and the cost-model item leave DCH-0. They are dispositioned in a new 2.1 "changes since the baseline" section rather than edited into the fixed main@de2f5f baseline, which stays immutable so later main movement does not rewrite a historical measurement. What #9960 does NOT resolve is kept open and sharpened. It made the state representable, declared and renderable; it read neither host back. Desired-versus- observed stays a live blocker, and the operator's two hand-edits make it sharper rather than weaker: a hand-set host value agreeing with a declared corpus value by coincidence is precisely what a silent carrier cannot distinguish from convergence, so the readback is the entire receipt. DCH-0 is renamed and its question narrowed to what actually remains: enrolment, and the live rendered-unit reconciliation. One defect found while verifying, reported and deliberately NOT repaired here because it is outside this lane: gunbc.spark.serving_desired still carries an earlier annotation asserting in the present tense that concurrency is not modeled in desired state, that the slot count divides the context window, and that the concurrency row is deliberately not smuggled in. All three are false, and they contradict that same module's own value and annotation about 140 lines below. A stale annotation is data the substrate cannot check, so nothing reds -- a 3 meaning fork inside a single authority. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * DCH: carry the open dispositions into the gate chain, and specify the tool surface instead of forward-referencing it Review 58352 (codex/gpt-5.6-sol) requested changes on two findings. Both are correct and both are repaired here rather than argued. FIRST: the review section identified five gate-changing questions and left them beside the gate chain rather than inside it, so DCH-0 could begin and complete while knowingly preserving the restart-mid-turn fail-open that Q3 names. Recording a class is not handling it, and 5 exists to answer these in advance. The gate chain now carries an admission rule -- a gate does not complete while a 5.2 question owning one of its dispositions is open -- with the bindings tabled and repeated in each gate. DCH-0 states its own ending explicitly: a drain or lease so a converge cannot land against an in-flight turn, or a 4b(3) row whose trigger names the drain capability. Never silence, and never a note that a restart is unlikely. The rule bars a gate's RECEIPT, not work inside it. SECOND: the section said the tool-surface quirks were "expanded there rather than here" while DCH-2 still listed four tool names. That is a promise naming a route that does not exist -- the same defect this plan refuses in others, committed by me in the act of writing the section that refuses it. DCH-2 now specifies the surface as six owed dispositions under one rule: a tool's contract is carried in its declared shape, and any limit it enforces is reported in its result rather than inferred from a truncated or missing one. Working directory, duration, input, truncation, edit matching, context exhaustion. Two of the relayed harness's choices are adopted deliberately because they are already the right shape -- a cap that announces its truncation, and an edit that refuses on zero or multiple matches rather than patching the first. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * DCH: carry the five returned dispositions into the gates, and add DCH-0r All five questions returned 2026-09-02. Recorded as rulings rather than positions; where a proposal was accepted with refinement, the refinement binds. The chain changes. Q3 rejected a 4b drop as the normal shape -- a drop reports a lower guarantee, it does not turn an unsafe arm green -- so the drain becomes its own gate rather than debt, and reordering removes the need for debt entirely: DCH-0 endpoint rows -> DCH-0r quiescent maintenance -> DCH-0c ComputeHost enrolment -> DCH-2 turn activation DCH-0r's mechanical blocker is verified in this tree rather than relayed, and verifying it found two refinements sharper than the finding as received. spark.serving_realization realizes EnableSystemUnit as enable, start AND restart in one effect, so restart has no independently matchable identity and no fence can guard it. First: that restart is UNCONDITIONAL, emitted whenever the effect is realized rather than when the definition changed, so converging the unit is destructive to in-flight work even when it changes nothing. Second, and the Sparks are on this path: the USER-unit closure already separates EnableUserUnit from StartUserUnit and carries NO restart effect at all -- so on the hosts this lane targets, a changed unit definition has no modeled route to take effect. That is a candidate cause for the observed-versus-desired drift DCH-0 must reconcile, and it is written as a candidate because nothing here has read a host back. Also carried: the admission fence and the drain are ONE protocol, because a bare active==0 read admits a turn immediately after it; and the guarantee grain is stated rather than assumed, since a census of DCH leases proves quiescence only for DCH-managed clients. DCH-2 gains an exit bar of four qualification groups and no longer exits on "four tools and one terminal turn". DCH-3 requires BOTH the qualification receipts and the unchanged 14-step terminal, neither substituting for the other. DCH-4 waits on that combined admission, because retiring a working runtime on an easy-path green is how a replacement erases a correctness distinction instead of completing. Q1 refines my boundary: whose failure is evidence, but the smallest closed causal scope containing the uncertainty decides the stop, and a supervisor alive in a typed stopped state is not an absorbing fallback. Q2 is accepted at typed- obligation grain and rejected at prose or character grain, with awaiting- verification deliberately not acceptance since the model's own EndTurn cannot grade the model's work. Q4 holds n=1 under survivorship to a single existential proposition that prior art does not establish for our harness at all. Q5 replaces both my options with two axes plus journal-before-effect ordering: a supervisor may move a record from running to interrupted and never to accepted. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * DCH: adopt the operator's only-and-default direction, and separate default from retirement Operator direction 2026-09-02: make the .dag harness the only and default harness, ignoring Claude and Codex entirely. Adopted, with the two halves kept apart because they have different preconditions. DEFAULT IS A SELECTION; RETIREMENT IS A DELETION. "Ignore the other providers" is discharged by making ours the default selection at the seam and stopping investment in the others. It does not require deleting them, and DCH-4's deletion still waits on DCH-3's receipt. Deleting them to satisfy "only" before that receipt would leave the repository with zero working harnesses -- the failure the replacement-migration doctrine calls erasing a correctness distinction rather than completing the replacement. The other variants stay frozen in the doctrine's sense until DCH-4. "ONLY" RAISES THE BAR RATHER THAN LOWERING IT. With three providers a weak DCH-3 is tolerable because a fallback exists; with one there is none, so every DCH-2 qualification class becomes load-bearing on the day it becomes default. The relayed transcript in 5 is exactly a record of what a sole harness with no fallback feels like when it breaks: three sessions dead, an exit status of zero, and an operator who cannot tell working from dead. The combined admission is what makes "only" survivable, so it is not tradeable for focus. Also added to the must-not-do list: do not delete or break the existing runtimes to satisfy "only" before DCH-3's receipt, or the count of working harnesses passes through zero. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * DCH: withdraw the RLM-first start barrier, and correct DCH-0r's identity model to three identities The reviewer superseded my freeze at ef8cf03 for a reason I should have caught myself: the plan still contained the OLD sequencing in two places -- the opening "RLM closes first; this lane starts after it" and the section 4 prohibition on starting before RLM's terminal receipt -- while DCH-3 had already been rewritten to the operator's 2026-09-02 direction. The document contradicted itself, and no freeze can make contradictory plan text admissible. Concurrent DCH and RLM is authorized. DCH-3 is a JOIN, not a start barrier: DCH-0 through DCH-2 build and qualify while RLM is blocked, and the default cutover waits for both sides. "Only" raises that bar rather than licensing an early one. DCH-0r's protocol carried two conflations and the correction is now the load-bearing part of the gate. THREE identities, not one: a stable service locus, a service-invocation identity that changes on every activation, and a running-definition identity stable for equal specs. The /proc environ stamp answers the THIRD -- it is invocation-bound evidence but not an invocation identity, because two successive restarts from one definition carry the same stamp and so it cannot prove the old process was replaced. And the fence is keyed by the LOCUS, not the incarnation. An incarnation-scoped fence stops blocking admissions at exactly the moment an unvalidated new incarnation appears mid-drain. The pre-restart invocation is a CAS condition and an evidence anchor, never the scope of the refusal. The unstamped reading is narrowed. "No stamp" establishes that the invocation cannot be shown to have started from the current desired definition; it does NOT establish that the process predates it, since a foreign or manually started process could have been launched later and would carry no stamp either. It still requires reactivation; chronology is just not what the evidence supports. Also added to the must-not-do list: do not land a change to an authority already inside another lane's frozen approved delta without serializing it. A clean textual composition still yields a blob nobody approved. dag/gunbc/fleet/fleet_converge_plan.dag is the live instance, shared with #9832. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 * DCH: model the behaviour-bearing serving package, add DCH-0p, and retire the degenerate-stop hypothesis to a mechanical cause Ruling 2026-09-02. The plan's model identity was weights-shaped and that is insufficient: same weight blob, same template, same runtime, same host, DIFFERENT EFFECTIVE STOP SET is materially different observable behaviour. So DCH-0b's subject becomes a ServingModelPackageIdentity carrying template, parser, parameters, the effective stop-sequence set, runtime release and decode realization, with a tag as an allocated handle and never the semantic identity. A package rebuilt from the same weights without a stop string is a DIFFERENT package; treating it as identical makes the repair invisible to convergence. New gate DCH-0p, after DCH-0b and DCH-0r and before DCH-2 qualification. It reads what the endpoint ACTUALLY serves rather than what a Modelfile or a tag says it serves, refuses with PackageConfigurationIncomplete rather than inferring a behaviour-bearing field, and admits a stop sequence only when it binds uniquely to a declared protocol boundary in that exact package's contract. A copied allowlist is not a role, and sentinel-looking values inherited from another lane are not our evidence. THE HYPOTHESIS THIS RETIRES IS THE REASON THE LANE EXISTS. 5 recorded, from the relayed transcript, that an assistant turn may stop at end_turn having announced work it did not perform -- one observation, survivorship-filtered, which I was careful to say established almost nothing. It now has a mechanical cause: a serving package carrying PARAMETER stop ### terminates generation on ordinary Markdown H3 output while returning a well-formed stream and stop_reason end_turn. Reproduced by a tiny prompt, removed by rebuilding the package. So the heading-then-stop symptom had a determinate external cause the whole time, and the hypothesis about model behaviour was a transport symptom read as a semantic one. It leaves the hypothesis list and becomes DCH-0p's subject. The streaming hypothesis is untouched and still has NO observation. DCH-2 gains two things. Its terminal semantics may never equate end_turn with natural completion, because at least two upstream causes produce that same label and the client cannot prove the native cause from a surface that erased it -- so DCH-0p and DCH-2 are two independent walls rather than one. And throughput becomes a construction rather than a threshold: a cache lineage joined to accounting that separates newly evaluated from reused input tokens, since a rate over total input divides mismatched subjects and yields a valid arithmetic operation that is not a hardware measurement. "No four-digit rate is reachable" is a view, which #9946 forbids as a correctness wall. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LSzWg5t7F22xEtbd83fzQ6 --------- Co-authored-by: gunbc-ci-auto-heal <gunbc-ci-auto-heal@users.noreply.github.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
The class
A decidable predicate is written over measured values — cost at or above N, latency under N, the rows a report printed — and its output is then treated as the population of the class under study. Membership in that set is a property of the measurement, not of the subject: it moves when the instrument moves, when load moves, when the reporter truncates, while the subject is unchanged. The set and the population are two objects joined by an assumption nobody states, and the assumption is false exactly when the measured quantity carries an unbounded term the threshold cannot see beneath.
What separates it from its neighbours is the route it arrives by
instrument_output_read_as_subject_contentis a report read past its own promise.censored_estimator_drops_its_own_tailis a statistic computed over survivors. Both are errors in the reading.This one is produced by diligence. The predicate is decidable, honestly derived, and very often the correct answer to a legitimate refusal — which is precisely why the next question does not get asked. Being careful is what makes it invisible: the diligence is spent on the repair, and the repair's correctness is what stops anyone asking whether it answers the question that was posed. A defect arriving inside a correct fix is not caught by the diligence that produced the fix.
Both receipts belong to the same author, three hours apart
That is why it is filed rather than blamed.
Repair, and a corollary worth carrying separately
Name the closed subject universe from the subject rather than from any measurement, and keep the threshold set — which is genuinely useful — under an honest name: an exposed attention subset for prioritising work.
The corollary, because conflating the two hides the seam: the evidence floor behind such a threshold may be monotone while the membership set is not. An observed floor only rises, so a constant derived from it only falls; individual identities still enter and leave the subset as their measured values vary. Monotonicity of the bound says nothing about stability of the set.
Recognition rule
When a set is produced by comparing a measured value against a threshold, ask what the set would contain if the measurement were taken again under different conditions — and ask it hardest when the predicate was just written to satisfy a reviewer, because that is the moment the answer feels settled.
Scope
Row plus roster entry in
gunbc.recurring_failure_mode, withDESIGN.mdanddocs/design-ledgers.mdregenerated by the narrowmain_wet_oneper path rather than hand-edited. Deliberately kept out of gunbc#9932 — that PR was re-authored twice under review, and a second class riding along there gets skimmed rather than read.