Skip to content

V4.1 Group A launch as a fleet mode (spark_v41_group_a_launch) - #12429

Merged
gunbai-bot[bot] merged 14 commits into
mainfrom
session/clever-gull-48-group-a-launch-wet
Sep 28, 2026
Merged

gunbai-bot[bot] merged 14 commits into
mainfrom
session/clever-gull-48-group-a-launch-wet

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

Makes the V4.1 capacity-measurement serve dispatchable as a fleet mode: mode=spark_v41_group_a_launch target=srv5. It serves at the measurement's shape: gpu-memory-utilization 0.80, 8192 batched tokens, max-num-seqs 16, 262144 context, prefix caching off, and one serve for the whole staircase.

Stacked on #12416 (production receipt row) and #12331. This diff shrinks once those land.

What the mode does

  1. Image readback on srv5-8. The launch uses a new held-image standing, V41HeldRuntimeStanding, not a produced one. V41RuntimeProduced is sole_constructor: only the build leg can mint it, and an image adopted under a tag refuses. So the launch never pretends it built the image. Instead it joins two facts:

    • the production receipt row v41_production_produced_image;
    • a readback on every host that the distribution's tag names exactly that digest (through vllm_image_target_read, the same read the distribution does).

    Outcomes:

    • wrong revision, or another image under the tag: refuses;
    • nothing under the tag: an obligation naming the distribution mode;
    • no readback taken: an obligation.
  2. Registry reading inside the digest. The build's own probe (v41_probe_image_in_place) already prints REGISTRY_COUNT plus one REGISTERED line per architecture. The launch runs it on the session host and folds the output into an ArchitectureRegistryReading. The fold is complete or nothing: the probe identity must join the receipt digest, and the count must equal the number of lines. No reading means the arm wall refuses with its own cause.

  3. Manifest. Derived exactly as the differential derives it: the row-store spec, then v41_engram_store_manifest over the published readings.

  4. Occupancy. Read on every host. The memory need is derived as the measurement's 0.80 share of gb10_unified_pool_gib.

  5. Admission. The existing v41_group_a_launch_admission against the live authority log. That log read is now one function shared by the dry run and the wet entry.

  6. Only if admitted:

    • Engram staging on each host, halting at the first failure;
    • then the arm transaction, spark_native_apply_transaction. It was extracted from Group B's spark_native_serving_apply_wet and is now shared by both entries: grant preflight, preserve, head-first apply, per-rank readback, commit or rollback.

Exit 0 means COMMITTED. The receipt (target/v41-group-a-launch-receipt.txt) carries every read whatever the verdict.

Evidence (in-session claim_batch)

  • v41_group_a_launch_witness_test: 23/23. New claims:

    • held on every host (positive control);
    • holding nothing is an obligation to distribute;
    • another image refuses;
    • no readback is an obligation;
    • joined, complete probe gives a reading;
    • unjoined, short or refused probe gives no reading;
    • the occupancy need.

    The largest claim costs 11,000 eval steps. The candidate runtime key and the manifest are supplied values: computing them reads wheels and stores, and they are other modules' subjects.

  • native_serving_roce_transport_witness_test: 19/19 (Group B apply refactor).

  • fleet-converge.yml regenerated with generated_artifact_gate main_wet_one. The diff is the new step and the mode lists only.

Open risks (not exercised)

  • Staging permissions. Engram staging writes /var/lib/gunbc/v41-engram-{manifest,touch} as the managed executor (RemoteFileAsSessionPrincipal). If that principal cannot write there, the run stops at staging with no unit touched, and the receipt says so.
  • Timeout. The step uses the 60-minute serving timeout.
  • Engram path under load. The staircase itself is the capacity measurement's wet runner (v41_wet_runner_frontier), not this PR.

🤖 Generated with Claude Code

Brian Searls and others added 9 commits September 26, 2026 02:20
…mport as an image capability fact

Arm A is the pre-#12304 image from run 36207135528 (config
sha256:ea39410e...). Arms are keyed on the image config digest every rank's
startup receipt carries, instead of a patch digest. Each rank's startup
reading records whether vllm._deepselect_C loaded, failed to import (a
recorded fact), or was not observed (refuses). 25/25 witness claims.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review 71462: the recorded DeepSelect import had no reader. The capacity
report now names the kernel its figures ran on; a failed import on any rank
reaches it with the logged line. 26/26 witness claims.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ckedReleased

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…he images serve

Arm A is EngramUpstreamPinned (images built before the design-B cutover serve
upstream DPShared/pinned storage); arm B is EngramFileBackedReleased (the first
image after the cutover). 28/28 witness claims.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ge binding, arm B

Fleet-converge run 36327364664 (srv8, the design-B cutover's patch set):
tag gunbc-vllm-dsv41-gb10:3cedb66f0ec0869e, config sha256:5de3c355...,
source_head d2d649e6, worktree_diff_sha256 e967c379...

- gunbc.spark.v41_runtime_image_converge v41_production_produced_image: the held
  receipt row; the distribution mode now moves this image (source srv8 and its
  digest from the row). The row re-derives its printed tag through the build's
  key fold (witness).
- gunbc.spark.v41_engram_differential_run v41_differential_image: bound to the
  production image, so spark_v41_engram_differential runs the overlay module
  from it.
- gunbc.spark.v41_capacity_measurement: arm B (EngramFileBackedReleased) bound
  to the production image; the binding frontier is dissolved (both arms bound).

9/9 distribution, 6/6 differential, 29/29 capacity measurement.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…gests

Arm A/B bindings and the differential image now name v41_arm_a_produced_image /
v41_production_produced_image; the wet-runner frontier names v41_engram_arm_of as
the fold it consumes (review 71799).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…registry probed in its digest, native arm transaction

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 27, 2026

Copy link
Copy Markdown
Contributor Author

About the 1048576 KiB-per-GiB note in review 71826: std.measure has no GiB-to-KiB conversion. The only route is through gibibyte_to_byte_size and a hand division by 1024, which is no less literal. I'm leaving it as is; if the note is worth acting on, the conversion belongs in std.measure as its own change. — sent from clever-gull-48

Brian Searls and others added 2 commits September 27, 2026 18:25
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…n-image' into session/clever-gull-48-group-a-launch-wet
@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Sep 27, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to a conflict with the base branch Sep 28, 2026
… fleet-converge.yml

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

The floor red at 74e5f19 is not caused by this change. Two wet witnesses failed:

  • spark_pair_serving_apply_wet_witness.the_apply_seam_reads_the_authority_before_planning_and_refuses_off_fleet
  • fabric_capacity_standing_wet_witness.the_live_entry_reads_the_authority_before_judging_and_refuses_off_fleet

Both assert that the runner cannot read the serving-authority log. On this PR's earlier heads they ran on srv3 runners and passed (runs 36336688221 and 36341124978). This run landed on srv1-06, where they failed (run 36370281310). The likeliest reason is that the log is readable from srv1, which the 'off fleet' premise rules out. So their verdict depends on which runner picks them up. I've re-run the failed jobs and flagged the witness premise to the owning lane. — sent from clever-gull-48

…yml from the .dag

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

About the fixed offset in v41_registry_count (review 72095): 15 is the length of the prefix "REGISTRY_COUNT ", and the slice only runs after starts_with has matched that exact prefix. A line in any other shape takes no slice. It is then refused by parse_int, and by the check that the count equals the number of REGISTERED lines. Deriving it as length(prefix) would be tidier, and I'll do that if this PR needs another push; on its own it isn't worth resetting CI on a head that's about to queue. — sent from clever-gull-48

@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Sep 28, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Sep 28, 2026
@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Sep 28, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to a conflict with the base branch Sep 28, 2026
…yml from the .dag

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Sep 28, 2026
Merged via the queue into main with commit 85672ee Sep 28, 2026
6 checks passed
@gunbai-bot
gunbai-bot Bot deleted the session/clever-gull-48-group-a-launch-wet branch September 28, 2026 10:13

Copy link
Copy Markdown
Contributor

Cross-lane handoff: GLM promotion now; common serving follow-ons here, not another model-specific stack

For the DS4.1 lane: Brian has directed promotion of GLM after the reported clean 262K×8 closing run, and explicitly requested that the subsequent work be coordinated with DS4.1 and generalized to retained past and future deployments. Full steering is recorded once at #12387 (comment) .

This merged PR and #12378 already provide the correct connection: parameterized plan_arm_launch_for, ArmContainerIdentity / ArmUnitIdentity, spark_native_arm_steps_of, and shared spark_native_apply_transaction. The GLM deployment branch at dee66ef133e7584a80e003c69826ec8074f43ec3 still carries the older per-unit apply implementation; its tested experiment cell and your main-side transaction must be reconciled, not copied into a third controller.

Please coordinate the shared follow-on ownership with the GLM lane: candidate/evidence applicability (using the existing identity grain split), per-rank/per-phase resource accounting, exclusive lifecycle/settlement, composed budgets, enforced ingress/queue policy, semantically correct metrics, and promotion/readback. GLM and DS4.1 should consume the same mechanisms and behavioral controls while providing different architecture/runtime/storage/protocol adapters. DS4.1's Engram, cache reconstruction and selected-kernel facts stay model/runtime facts, not generic assumptions or GLM constants.

Immediate boundary to preserve: GLM's clean churn used a client-local flock to serialize large requests. It must be an actual enforced production restriction before that constrained qualification can support the exposed route. Its sampled process-footprint gap is not a fragmentation bound. These are interface obligations shared with DS4.1, not reasons for another wholesale GLM cold-work campaign.

Generalization completion must include existing callers, not just two new generic types: migrate retained older model routes or explicitly retire them, preserve historic receipts without upgrading their claims, and require future model deployments to enter through the same conformance path. Do not hold GLM promotion for this broader migration or W1b optimization; do not interpret this message as authorization to deploy DS4.1 or override either lane's existing access/admission gates.

This is a handoff comment on the latest specific DS4.1 launch integration I found, not a new review of this already-merged PR. No live operation performed.

@briansrls
briansrls restored the session/clever-gull-48-group-a-launch-wet branch September 29, 2026 19:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant