Repository navigation
Model DSpark as its own authority, free srv6 from training, and make the fabric desired state - #10009
Conversation
…the fabric desired state Three changes, one objective: make fleet convergence safe to run on the Sparks. DSPARK IS DEEPSEEK'S, NOT NVIDIA'S. "dspark" was used across two sessions and an operator conversation to mean NVIDIA's dgx-spark-playbooks; the operator meant DSpark, DeepSeek's speculative-decoding algorithm. One spelling, two independently governed upstreams -- the DESIGN section 3 meaning fork, found only when the operator asked directly. extdeps.deepseek.deepspec is the DeepSeek authority, verified against the DeepSpec repository and the Hub rather than recalled. Its safety content is that a draft is BOUND TO ONE TARGET. Speculative decoding is output-equivalent by design, so a draft paired with the wrong target does not produce wrong text -- it produces the same text more slowly, while every throughput number measured from it describes a configuration nobody intended. The binding carries its target and there is no constructor for an unbound draft. The attachment SHAPE is a second fact, and modelling it as one was a defect. A draft ships either as a separate artifact (the antirez GGUFs) or fused into the checkpoint (deepseek-ai/DeepSeek-V4-Flash-DSpark: 48 shards and 166.9 GB against the plain repository's 46 and 159.6 GB). In the fused shape there is no download that yields the base alone, so a consumer sizing from base_total plans a residency no disk will hold. SRV6 IS NO LONGER RESERVED FOR TRAINING, by operator decision: role is converged desired state, not a dedication. Consequence in serving_converge_plan: srv6 receives full serving desired members rather than baseline. Three witness claims that transcribed the current roster are rewritten as the invariants they were reaching for -- disjointness, and the desired-hosts/serving-cells bijection -- so a role change moves the roster without a test going red. The role lookup and the role-scoped planner now take the roster as a parameter, because an arm's evidence must not depend on which machines are currently assigned what. The production-root retirement claims that needed a live training cell are removed under a declared section 4b(3) rung drop. THE FABRIC IS DESIRED STATE BECAUSE RUNTIME CONFIGURATION WAS LOST. The RoCE addressing on srv5/srv6 was applied by hand, verified at 200 Gb/s with jumbo frames, reported as configured, and destroyed by one power cycle: nmcli writes land in /run/NetworkManager, which is tmpfs, and the durable artifact is the generated /etc/netplan YAML. gunbc.spark.fabric_link_desired owns that, read back off both hosts rather than authored from intent, and a rail whose ends name different interfaces has no constructor -- the crossed configuration passed a jumbo ping and was still wrong, which is a failure that tests green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
Two review findings, both correct. MTU was `mtu: Int` on FabricEndpoint and `data fabric_mtu_bytes: Int = 9000` -- a byte quantity as a bare scalar, in a record whose network half had just been given PrefixLength for exactly that reason. It carries std.measure's ByteSize now. The scale is what a 9000-vs-1500 mismatch turns on, and that mismatch is one of the constructor's five refusal arms. The section 4b(3) drop was declared as a block comment in the witness file, and DESIGN section 4b says declared drops are rostered in full in docs/design-ledgers.md under authority gunbc.rung_drop. An in-source comment does not satisfy 'no untracked stall' -- it is not countable and not prioritizable. The row is now spark_role_scoped_retirement_production_root in gunbc.rung_drop, with previous rung, temporary rung, reason, bounded population and a restoration trigger that names the capability; the witness comment is reduced to a pointer, because a second copy of a declaration is a second authority for it. docs/design-ledgers.md and DESIGN.md are the generated-artifact actuator's own output (main_wet_one), not hand-edited. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
…1-cell-role-converged # Conflicts: # DESIGN.md # dag/gunbc/rung_drop.dag # docs/design-ledgers.md
Operator-priority ruling: safe fleet-converge E2E is this PR or its immediate successorThe operator’s controlling objective is now explicit: stop hand-editing Sparks and prove a safe fleet-converge E2E as soon as possible — in this PR if the composition remains reviewable, otherwise in the immediately following PR with no intervening Spark feature work. This PR correctly removes the premature permanent training dedication from
Required first terminal, on exactly one restored
Compose rather than rebuild: #9897 supplies exact evidence identity; #9984 supplies typed reactivation vocabulary; #9986 supplies running-definition/invocation observation; the existing Operator decision carried: Tracking/terminal: #10001. |
|
CI investigated. Neither failing check is attributable to this PR, and I am not pushing a speculative fix over them. Evidence below so it can be checked rather than taken.
This branch changes no Rust whatsoever — This branch's own Same class on the earlier head. The build lane on
The diagnostic distinguishes the two cases itself and says which one this is: the toolchain was disturbed mid-run. And on #9897 the floor refused with So there are two independent things, and neither is this change: a genuine pre-existing red on What I verified locally instead, since the executing evidence is what matters:
Merge readiness is therefore blocked on infrastructure and on a — sent from eager-pike-541 |
Engine-pivot ruling — exact runtime scope, and no change to the first safe-converge terminal
Head/expert divisibility is likewise strategy-specific. It may be a constructor condition for an exact tensor-parallel plan whose implementation requires even partitioning, but it is not a universal assembly law and must not be applied to llama.cpp RPC. RPC exposes remote devices and distributes model tensors/KV by its own rules; a three-host RPC candidate must be judged by that implementation, not by TP=3 arithmetic.
The 155 GiB safetensor copies are therefore |
Queue-stall boundary: off-tree RPC probing is allowed, but it does not interposeThe existing “no intervening Spark work” ruling constrains PR/authority order. During the shared-runner availability refusal, a quarantined llama.cpp RPC experiment may run without changing this PR or the immediate #10001 terminal. It must remain off-tree and non-authoritative: no desired-state or role change, no PR, no evidence transport from Ollama, no claim of P0a/P0b completion, and no use of the result to delay the exact single-host Ollama convergence receipt. The probe must use exact runtime/artifact/host/configuration identities in its raw receipt, private-link RPC exposure, attempt-owned transient processes, and positive teardown/re-observation of the managed Ollama state. Any need to stop/rewrite/restart the managed unit or any host instability ends the probe rather than broadening #10009. Publication order remains: #10009 on its accepted subject, then #10001 as the immediate Spark-serving successor unless the full wet terminal is absorbed here. |
…1-cell-role-converged # Conflicts: # DESIGN.md # docs/design-ledgers.md
Four claims in v2.test.claim.spark_observation_scope went red when srv6's training dedication was withdrawn -- they asserted the LIVE role counts (length(training) == 1, length(spark_serving_cell_hosts()) == 1) against the production assignments. No behaviour changed; the machines did. That is the same defect spark_cell_role_retirement_witness_test was repaired for, and cell_role's own note already states the principle: a capability's evidence must not depend on which machines are currently assigned what. spark_cell_role_in carried that principle for a single host's lookup. It stopped there, so every claim about a ROLE POPULATION kept the dependency. spark_hosts_with_role_in lifts it to the population, and spark_hosts_with_role becomes its application over the production roster. The four claims now read a fixture roster (srv5 serving, srv6 training) and keep their subject: the observation scope is PHYSICAL and wider than either role roster, so scoping observation by a role-scoped roster loses whatever the other role holds. Each remains discriminating -- shrink the scope to one host and all four go red. Verified by execution: the four return true at this head, and returned false at 0c4ed0c where the required floor reported failed=4. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G
New live divergence from the exploratory RPC run: adopt here or revert before #10001The RPC probe created durable fabric state on The files are If the operator elects to keep the new links, this PR is the correct and only permitted absorption point before SAFE-SPARK-CONVERGE-0: add the exact srv5/srv6 links through the existing If the operator declines adoption, revert the persistent profiles and positively observe absence. In either arm, #10001 must begin from a fresh post-disposition observation of both hosts. Undeclared durable state may not be treated as harmless merely because the addresses improve an exploratory runtime. This is not authorization to add llama.cpp RPC runtime/deployment authority here. The RPC measurements remain quarantined; #10009's subject is the durable fabric and role state, while the first production terminal remains the exact single-host Ollama converge. |
Sequencing correction: do not reopen this green head solely for rails created after itThe post-head RPC run created durable Therefore, if the operator accepts the durable links, #10009 may land on its current accepted subject without being moved solely to add them. Put the exact If the operator declines the links, the alternative is explicit reversion plus observed absence. Until the operator selects one arm, make no further network mutation. This amendment does not admit standalone llama.cpp RPC authority, transport its exploratory evidence, or move the first terminal. Publication remains: Separate correction: SHA-256 |
Landing caveat: rail sequencing does not waive current-main compositionThe comment above removes the new rails as a reason to reopen this exact head. It is not merge authorization for a stale composition.
Consequently, if the operator selects Adopt before that mandatory recomposition, it is economical and admissible to add the exact |
…1-cell-role-converged # Conflicts: # DESIGN.md # dag/gunbc/rung_drop.dag # docs/design-ledgers.md
Three changes with one objective: make fleet convergence safe to run on the Sparks. The operator asked for this "ASAP, this or next PR", and separately for the correct DSpark to be modeled in
extdepsand anchored in convergence.1. DSpark is DeepSeek's, not NVIDIA's — a §3 meaning fork found by the operator
"dspark" was used across two sessions and an operator conversation to mean NVIDIA's
dgx-spark-playbooks(the DGX Spark hardware recipes,connect-two-sparksand its vLLM multi-node section). The operator meant DSpark, DeepSeek's speculative-decoding algorithm. One spelling, two independently governed upstreams, nothing in common. A fabric, a Ray topology and a privileged-container decision were all pursued under the wrong referent before anyone asked.extdeps.deepseek.deepspecis the DeepSeek authority. Every fact in it was verified againstdeepseek-ai/DeepSpecand the Hub rather than recalled: the three algorithms DeepSpec implements (DSpark, DFlash, Eagle3) with their papers as further citations, and the released(algorithm × target)checkpoint grid.Its safety content is that a draft is bound to one target. Speculative decoding is output-equivalent by design — the draft proposes, the target verifies, rejected tokens are discarded. So a draft paired with the wrong target does not produce wrong text; it produces the same text, more slowly, while every acceptance-rate and throughput number measured from it describes a configuration nobody intended. A performance claim that silently detaches from its subject. The binding carries its target, and there is no constructor for an unbound draft.
The attachment shape is a second fact, and modelling it as one was a defect. A draft ships either as a separate artifact (the three
*-DSpark-support.gguffiles in the antirez distribution) or fused into the checkpoint:deepseek-ai/DeepSeek-V4-Flash-DSparkis 48 shards and 166.9 GB against the plain repository's 46 and 159.6 GB — same architecture, same fp8/e4m3/ue8m0 quantization, plusnum_nextn_predict_layers: 1anddspark_block_size: 5. DeepSeek's own model card: "not a new model. It is the same checkpoint with an additional speculative decoding module attached." In the fused shape there is no download that yields the base alone, so a consumer sizing frombase_totalplans for a residency no disk will ever hold, and 159.6 GB names a repository deliberately not in use.Also grounded here, because it was being carried as an open project risk: vLLM does implement DSpark, via
--speculative-config '{"method":"dspark",...}'. What remains unverified is whether a specific engine build carries it — an observation about that build, not a fact about this upstream.2. srv6 is no longer reserved for training
Operator decision, recorded with its basis: "i wouldn't dedicate a whole node for any task - just have it converge and serve whatever task we converge it to." Role is converged desired state, not a dedication. Training remains a modeled capability in
gunbc.spark.training_ready; it stops being a node reservation. Consequence inspark_serving_role_scoped_desired_members: srv6 receives full serving desired members rather than baseline-only.Three witness claims transcribed the current roster (
length(serving) == 1 && srv5 serves && srv6 trains) and one said in its own comment that it goes red the moment srv6 reverts. That turned an operator decision into a failing test. They are rewritten as the invariants they were reaching for — disjointness, and the desired-hosts/serving-cells bijection — which mention no host and survive any roster change. DESIGN §5: completeness is an identity join, not a count equality.spark_cell_role_inandspark_serving_role_scoped_desired_members_innow take the assignment roster as a parameter, so the training arm stays exercisable regardless of which machines currently hold which role. The production-root retirement claims that needed a live training cell to have a subject are removed under a declared §4b(3) rung drop, with previous rung, temporary rung, reason, bounded population, and a restoration trigger that names the capability (roster threading through the plan artifact) rather than "put a training cell back" — which would re-create the very coupling the drop records.3. The fabric is desired state, because runtime configuration was lost
The 200 Gb/s RoCE addressing on srv5/srv6 was applied at runtime by hand, verified working at MTU 9000 with a clean 8972-byte jumbo ping, reported by me as configured — and destroyed entirely by one power cycle. I reported a runtime observation as durable state.
The mechanism: these hosts run netplan with
renderer: NetworkManager, and annmcliwrite is intercepted by the netplan-NM integration into/run/NetworkManager/system-connections/, which is tmpfs. The durable artifact is the generated/etc/netplan/90-NM-<uuid>.yaml, and only that. So a convergence that actuatesnmcliconverges a host into a state that does not survive its next boot, while reporting success.gunbc.spark.fabric_link_desiredowns the netplan facts, read back off both hosts rather than authored from intent. A rail whose two ends name different interfaces has no constructor: the crossed configuration existed, was applied by hand, passed a jumbo ping, and was still wrong, because NCCL selects a rail by matching interface to subnet. That is a failure that tests green, which is why it is a constructor condition rather than a diagnostic.extdeps.network.ipv4gainsPrefixLength,render_ipv4_cidrandipv4_same_network— the CIDR spelling every router table and netplan file writes, and a same-network predicate that refuses for a prefix it cannot answer exactly rather than assuming a match.srv7/srv8 are named, not enrolled. That module's own note makes the distinction structural: membership in
endpointsis enrollment. Adoption is five separable acts and this is the second alone.Evidence
All executed, all green:
w_all_spark_fabric_claims_hold— new witness. Positive control plus five refusal arms (same host, crossed interfaces, unrelated subnets, identical addresses, MTU mismatch), the subnet-boundary discriminator, and the unanswerable-prefix refusal.w_the_planner_scopes_desired_state_by_role,w_the_production_roster_reaches_the_serving_arm,w_spark_roles_are_disjoint_over_the_assignment,w_serving_desired_hosts_are_the_serving_cells,w_serving_cell_is_not_offerable_to_the_training_fabricdeepspec_released_dspark_bindingsreturns all four bindings with the correct attachment shapes and block sizes (7 for the Qwen grid, 5 for V4-Flash — block size is a property of the trained draft, not of the algorithm, which is why it is a field).What this does not do
It does not yet repoint
gunbc.spark.serving_desiredoffgpt-oss:120b, and it does not yet consume #9897's memory-fit function inserving_converge_plan. Those are the remaining two blockers before the wet slice can run, and they land next — the aggregate-memory fit across a two-host assembly is what the operator's "~160 GB across two Sparks, rest for slots" actually needs, and it depends on #9897 landing first.🤖 Generated with Claude Code
https://claude.ai/code/session_01U397y4s3dSBof7vGPAX87G