Repository navigation
Gap analysis: hand-edited Spark state with no modeled authority - #9991
Conversation
…ebuild Every divergence read off the live hosts against the module named as its authority, not recalled. The load-bearing finding is that convergence is armed and wrong in three directions at once: modeled context is 131072 against a live 1048576, the served model is gpt-oss against a live DeepSeek-V4, and OLLAMA_NUM_PARALLEL appears NOWHERE in the corpus -- so convergence does not overwrite the slot count, it deletes it. None of those refuse. The unit rewrites, the service restarts, and it serves, narrower and serialized, with every number the serving decision is reasoned from silently invalidated. Also rostered: the second Spark pair is provisioned and serving while unknown to cell_role, the fleet slot roster and network identity subsumption; residency is unmodeled on both sides so a 94 GB model evicts and reloads unmeasured; and credentials were hand-carried past a modeled SecretRef path that already exists. One finding turns back on my own PR: the realization carrier keys a runner on argv, and every runner here is bare `ollama serve` with its whole configuration in Environment= lines. On this fleet that identity discriminates almost nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
The context and slot-count divergences this document led with were false. Both were derived by running the renderer from a session branch based on 373b8d1 while main had advanced: main desires 1048576 context, carries the slot count as a typed PositiveSlotCount of 4 through extdeps.ollama.server_env, renders it, and witnesses the rendering. The freeze recommendation that followed from them is withdrawn. The error is its own class and the document now says so: a real instrument was run, a real number came back, and it was read as a fact about the fleet when it was a fact about my branch. A divergence report is only as current as the tree it was measured from. What survives: the served model (desired gpt-oss, live DeepSeek-V4), residency unmodeled on both sides, the second pair outside every authority, and hand-carried credentials past a modeled SecretRef. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
|
CI on this branch is red for a reason that is not in this diff, and I want that on the record rather than silently retried. This PR is one new markdown file, +145 lines, nothing else ( The failing lane is That same failure is present on main itself at So this is inherited, not caused. I am deliberately not regenerating the artifact on a docs-only branch: that would put a generated-surface change in a PR whose entire subject is a planning document, and the regeneration belongs wherever the drift was introduced. Main has since advanced past If the drift turns out to persist on current main, that is a main-side repair and I will raise it separately rather than smuggle it through here. — sent from eager-pike-541 |
|
Operational supersession: #10001 is now the controlling immediate path to safe Spark fleet convergence. Before this document can govern operations, its stale PR body and repeated paragraph must be corrected: context and Do not broaden #9991 into implementation. Link the corrected gap record to #10001, where completion requires one exact-host wet plan→apply→full readback→zero replan→terminal health transaction plus a second zero-effect noop. |
…rect the second pair Two divergences this document reported have moved, and leaving either as written would make the document assert a state the tree no longer has. SERVED MODEL IS CLOSED. serving_desired names the DeepSeek manifest through gunbc.ollama_model_resolution, admission requires an exact runtime version and a manifest-identity match rather than a name match, and the stale `Description=` spelling is now derived from the model ref instead of authored beside it. What did not close -- reproducing the local artifact from desired state -- is a declared rung drop rather than silence, and the section says so rather than reporting the whole axis as done. THE SECOND PAIR IS NOT TWO SERVING HOSTS. The section described two independent ollama hosts; .232 is a llama-server front door whose weights sit on .233 behind ggml-rpc-server over the RoCE fabric. That is one serving unit split across two machines, and the correction matters because admitting them as two serving cells would plan an 86.7 GB load onto a peer with ~16 GiB free and swap disabled -- an OOM kill that takes the front door with it. The membership gap is unchanged; what membership would have to MEAN changed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
The operator asked for an accumulation of what has been hand-edited on the live Sparks while the fleet is assembled by hand.
The reason it is worth a document rather than a plan line is one finding: a convergence run today would degrade the fleet without failing. The modeled desired unit differs from the live unit on BOTH serving hosts in three directions at once, all pointing the same way:
131072modeled against1048576live -- below the 400k floorgunbc.model.choiceexists to defendgpt-ossmodeled againsthf.co/antirez/deepseek-v4-ggufliveOLLAMA_NUM_PARALLELappears nowhere in the corpus, so convergence does not overwrite the slot count, it DELETES it and the runner falls back to a defaultNothing on that path refuses. The unit rewrites, the service restarts, it serves -- narrower and serialized.
Everything was read off the live hosts and compared against the module named as its authority. Desired side:
gunbc.spark.serving_unit_render spark_serving_desired_user_unit_text. Live side: the unit file and/api/psper host.Also rostered: the second pair (
spark-c2b1,spark-ac79) is serving while unknown tocell_role, the fleet slot roster, andnetwork_identity_subsumption; residency is unmodeled on both sides; credentials were hand-carried pastgunbc.spark.credential_workflow.One row turns back on my own #9897: it keys realization identity on argv, and every runner here is bare
ollama servewith its configuration inEnvironment=lines -- so on this fleet that identity discriminates almost nothing. Widening it needs a ruling on which variables are configuration and which are ambient.Doc only -- no code, no substrate.
🤖 Generated with Claude Code
https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G