Repository navigation
V4.1 Group A launch as a fleet mode (spark_v41_group_a_launch) - #12429
Conversation
…mport as an image capability fact Arm A is the pre-#12304 image from run 36207135528 (config sha256:ea39410e...). Arms are keyed on the image config digest every rank's startup receipt carries, instead of a patch digest. Each rank's startup reading records whether vllm._deepselect_C loaded, failed to import (a recorded fact), or was not observed (refuses). 25/25 witness claims. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review 71462: the recorded DeepSelect import had no reader. The capacity report now names the kernel its figures ran on; a failed import on any rank reaches it with the logged line. 26/26 witness claims. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ckedReleased Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…he images serve Arm A is EngramUpstreamPinned (images built before the design-B cutover serve upstream DPShared/pinned storage); arm B is EngramFileBackedReleased (the first image after the cutover). 28/28 witness claims. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ge binding, arm B Fleet-converge run 36327364664 (srv8, the design-B cutover's patch set): tag gunbc-vllm-dsv41-gb10:3cedb66f0ec0869e, config sha256:5de3c355..., source_head d2d649e6, worktree_diff_sha256 e967c379... - gunbc.spark.v41_runtime_image_converge v41_production_produced_image: the held receipt row; the distribution mode now moves this image (source srv8 and its digest from the row). The row re-derives its printed tag through the build's key fold (witness). - gunbc.spark.v41_engram_differential_run v41_differential_image: bound to the production image, so spark_v41_engram_differential runs the overlay module from it. - gunbc.spark.v41_capacity_measurement: arm B (EngramFileBackedReleased) bound to the production image; the binding frontier is dissolved (both arms bound). 9/9 distribution, 6/6 differential, 29/29 capacity measurement. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…gests Arm A/B bindings and the differential image now name v41_arm_a_produced_image / v41_production_produced_image; the wet-runner frontier names v41_engram_arm_of as the fold it consumes (review 71799). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…registry probed in its digest, native arm transaction Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
About the 1048576 KiB-per-GiB note in review 71826: |
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…n-image' into session/clever-gull-48-group-a-launch-wet
… fleet-converge.yml Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
The floor red at 74e5f19 is not caused by this change. Two wet witnesses failed:
Both assert that the runner cannot read the serving-authority log. On this PR's earlier heads they ran on srv3 runners and passed (runs 36336688221 and 36341124978). This run landed on srv1-06, where they failed (run 36370281310). The likeliest reason is that the log is readable from srv1, which the 'off fleet' premise rules out. So their verdict depends on which runner picks them up. I've re-run the failed jobs and flagged the witness premise to the owning lane. — sent from clever-gull-48 |
…yml from the .dag Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
About the fixed offset in |
…yml from the .dag Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Cross-lane handoff: GLM promotion now; common serving follow-ons here, not another model-specific stackFor the DS4.1 lane: Brian has directed promotion of GLM after the reported clean 262K×8 closing run, and explicitly requested that the subsequent work be coordinated with DS4.1 and generalized to retained past and future deployments. Full steering is recorded once at #12387 (comment) . This merged PR and #12378 already provide the correct connection: parameterized Please coordinate the shared follow-on ownership with the GLM lane: candidate/evidence applicability (using the existing identity grain split), per-rank/per-phase resource accounting, exclusive lifecycle/settlement, composed budgets, enforced ingress/queue policy, semantically correct metrics, and promotion/readback. GLM and DS4.1 should consume the same mechanisms and behavioral controls while providing different architecture/runtime/storage/protocol adapters. DS4.1's Engram, cache reconstruction and selected-kernel facts stay model/runtime facts, not generic assumptions or GLM constants. Immediate boundary to preserve: GLM's clean churn used a client-local Generalization completion must include existing callers, not just two new generic types: migrate retained older model routes or explicitly retire them, preserve historic receipts without upgrading their claims, and require future model deployments to enter through the same conformance path. Do not hold GLM promotion for this broader migration or W1b optimization; do not interpret this message as authorization to deploy DS4.1 or override either lane's existing access/admission gates. This is a handoff comment on the latest specific DS4.1 launch integration I found, not a new review of this already-merged PR. No live operation performed. |
Makes the V4.1 capacity-measurement serve dispatchable as a fleet mode:
mode=spark_v41_group_a_launch target=srv5. It serves at the measurement's shape: gpu-memory-utilization 0.80, 8192 batched tokens, max-num-seqs 16, 262144 context, prefix caching off, and one serve for the whole staircase.Stacked on #12416 (production receipt row) and #12331. This diff shrinks once those land.
What the mode does
Image readback on srv5-8. The launch uses a new held-image standing,
V41HeldRuntimeStanding, not a produced one.V41RuntimeProducedis sole_constructor: only the build leg can mint it, and an image adopted under a tag refuses. So the launch never pretends it built the image. Instead it joins two facts:v41_production_produced_image;vllm_image_target_read, the same read the distribution does).Outcomes:
Registry reading inside the digest. The build's own probe (
v41_probe_image_in_place) already printsREGISTRY_COUNTplus oneREGISTEREDline per architecture. The launch runs it on the session host and folds the output into anArchitectureRegistryReading. The fold is complete or nothing: the probe identity must join the receipt digest, and the count must equal the number of lines. No reading means the arm wall refuses with its own cause.Manifest. Derived exactly as the differential derives it: the row-store spec, then
v41_engram_store_manifestover the published readings.Occupancy. Read on every host. The memory need is derived as the measurement's 0.80 share of
gb10_unified_pool_gib.Admission. The existing
v41_group_a_launch_admissionagainst the live authority log. That log read is now one function shared by the dry run and the wet entry.Only if admitted:
spark_native_apply_transaction. It was extracted from Group B'sspark_native_serving_apply_wetand is now shared by both entries: grant preflight, preserve, head-first apply, per-rank readback, commit or rollback.Exit 0 means COMMITTED. The receipt (
target/v41-group-a-launch-receipt.txt) carries every read whatever the verdict.Evidence (in-session claim_batch)
v41_group_a_launch_witness_test: 23/23. New claims:The largest claim costs 11,000 eval steps. The candidate runtime key and the manifest are supplied values: computing them reads wheels and stores, and they are other modules' subjects.
native_serving_roce_transport_witness_test: 19/19 (Group B apply refactor).fleet-converge.ymlregenerated withgenerated_artifact_gate main_wet_one. The diff is the new step and the mode lists only.Open risks (not exercised)
/var/lib/gunbc/v41-engram-{manifest,touch}as the managed executor (RemoteFileAsSessionPrincipal). If that principal cannot write there, the run stops at staging with no unit touched, and the receipt says so.v41_wet_runner_frontier), not this PR.🤖 Generated with Claude Code