Repository navigation
V4.1 TP4 launch on Group A (srv5-8): a modeled dry run through the mp arm launcher - #12378
Conversation
…ich stops being GLM-only The V4.1 image (upstream vLLM d2d649e6's vllm-openai target) ships no Ray, so it cannot run on the V4 Ray pair. Per proud-deer-538's ruling, V4.1 is instead a new call site of the existing mp-backend multi-host arm launcher (serving_arm_launch / native_serving_realization / native_serving_apply). Changes: - The launcher's GLM literals move to their owners: ArmContainerIdentity (container name, rendezvous port, shm size, extra mounts), CheckpointVllmParsers on the checkpoint row, and ArmUnitIdentity for unit names and logs. - Group B renders byte-identical (57,535 bytes compared before and after). - extdeps.vllm.engine_args VllmEngramConfigRequest adds vLLM's EngramConfig in its dotted argv form. - serving_arm gains a V4.1 checkpoint row (backbone weight), a produced-runtime variant carrying its admitted identity, and deepseek_v41_measurement_profile. - gunbc.spark.v41_group_a_launch plans all four hosts with distinct names, the Engram env and mounts, and per-host staging. It adds an admission fold that reads the authority log (successor may launch, candidate key, image key, held hosts), and a dry run that renders every host's create argv, unit and scripts. - Inputs from open PRs (#12351, #12359, #12371, the produced image) are typed obligations that make the dry run refuse by name. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…llowance, not zero On a GB10 the page cache is the same unified pool as the weights. So the V4.1 checkpoint row now declares resident_beyond_weights = PageCacheFromNonKvHeadroom. The fit wall charges it as the pool share the launch profile's memory fraction leaves outside the engine, which is the axis the capacity measurement set for exactly this. It is not zero, and it is not the whole ~203 GB table. The value is a bet: v41_engram_resident_bets, judged against v41_after_requests_resident, confirms or falsifies it. Without a profile the charge is unknown, and serving_arm_admits refuses rather than fitting on weights alone. The launch planner asks serving_arm_admits_under with its profile. GLM rows declare NoResidentBeyondWeights, and Group B still renders byte-identical. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…t fns Review 71644: the V4.1 Engram staging was a hand-assembled shell body. It is now three gunbc.remote_operation PlannedRemoteSteps: - the extdeps.tools.mkdir argv for the manifest and receipt directories; - a gunbc.typed_remote_file_write sealed to the manifest directory. The receipt directory needs no clearing, because the overlay writes each rank's receipt by temp file and rename. Required floor: test fns may not reference test code. Deleted: - the aggregators in v41_group_a_launch_witness and glm_group_b_launch_authority_witness (refused by the floor); - serving_arm_launch_witness's aggregator, together with its eight admitted test-reference debt rows in v1.compile (debt paid). The v1_compiler_compile.rs mirror is regenerated; it differs only by those eight rows. Every claim still runs on its own. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Addressing review 71644 (the Engram staging was a hand-built shell script) in 7a35b1d, taking the reviewer's preferred fix rather than a marker: the staging is no longer a shell body.
Two lines of the old script were dropped rather than modeled, because they were redundant:
The dry-run witness now asserts the manifest appears as a modeled — sent from zesty-ant-800 |
… MQ-1: a let value reads through body_lower_value_read beside main's carried annotation; the retired chunks stay retired and main's emitted_add chunk is kept; ledger receipts kept
DS4.1 critical path (work item adhoc-894b491d-00e): a modeled launch of the V4.1 four-rank (TP4) arm on Group A (srv5–8), as a dry run. Nothing in this PR launches on a real host. It has no wet entry; the only live read is the authority log. The design was reported to proud-deer-538 and approved before building, including one ruled change of route (below).
The route: Group B's mp-backend arm launcher, not the Ray pair
The brief asked to re-derive from the V4 Ray pair realization. That cannot run V4.1: the V4.1 image is upstream vLLM d2d649e6's
vllm-openaitarget, and neither that revision'srequirements/*.txtnor itsdocker/Dockerfileinstalls Ray. Ruled by proud-deer-538 (option A): use the launcher that already fits, a group-parameterized, Ray-free multi-host launcher (--distributed-executor-backend mp,--nnodes/--node-rank/--master-addrover the fabric rail, headless followers,FourSparkTp4):gunbc.spark.serving_arm_launchplan_arm_launch_forplans it.gunbc.spark.native_serving_realizationrenders the units.gunbc.spark.native_serving_applyis the transactional apply (preflight / preserve / head-first apply / per-rank readback / commit or rollback).Ray is not added to the V4.1 image, because that would change the image key. There is still one realization, parameterized (DESIGN §3); V4.1 is a new call site, like
gunbc.spark.glm_group_b_launch.(The brief also said the pair is TP2 over two hosts with its head on srv7. On main the pair is already TP = |Group A| = 4, with its head
spark_pair_serving_head(FabricGroupA).)What changes
The launcher stops being a GLM launcher. Each GLM literal it spelled moves to the authority that owns it:
ArmContainerIdentity, a value the caller names.glm_native_arm_containercarries Group B's unchanged values.CheckpointVllmParsers). Which parser decodes a model's output is a fact about the model. An optionaltokenizer_moderenders no flag when absent.ArmUnitIdentity.glm_native_arm_unitscarries the GLM ones.native_serving_apply, each step carries its own log path and front door, and "head" isnode_rank == 0instead of a comparison against GLM's head-unit name.spark_native_arm_steps_ofplans any arm. The stage scripts take the container name.Upstream Engram request (
extdeps.vllm.engine_argsVllmEngramConfigRequest): vLLM d2d649e6'sEngramConfig, rendered in upstream's own dotted form--engram-config.cpu_offload true --engram-config.dp_shared_memory false.vllm/utils/argparse_utils.pyrewrites the dotted form into the nested JSON, so the words are plain and fit a systemd Exec line. OnServingLaunchProfileit enters the profile key only when explicit, so every existing GLM subject key is unchanged.V4.1 rows in
gunbc.spark.serving_arm:DeepseekV41FlashDba1be0acheckpoint row. Placement is wherev41_checkpoint_materializeinstalls the checkpoint; both now read onespark_local_checkpoint_directory. The parsers and tokenizer mode are the vLLM recipe's, which binds that recipe's "DSV41-1 launch argv binds …" frontier rows.resident_beyond_weights: PageCacheFromNonKvHeadroom. The fit wall charges it as the pool share the profile's--gpu-memory-utilizationleaves outside the engine. That is the axis the measurement set for exactly this: at 0.80 it is 97 GiB across four ranks (about 24 GiB per rank). It is not zero, and it is not the whole table.v41_capacity_measurementv41_engram_resident_betsagainstv41_after_requests_residentconfirms or falsifies it. Note: the stress model's bets (19/30/37/46 GiB per rank at x2/x4/x6/x16) exceed 24 GiB from x4 up, so the staircase's refault/PSI stop rules are what will discover whether that pressure matters.serving_arm_admitsrefuses the row rather than fitting it on weights alone. The planner usesserving_arm_admits_underwith its profile.GunbcDeepseekV41EngramProduced { artifact, registry_reading }runtime variant. The V4.1 image is produced by the lane and admitted at run time, so its identity is a value, not a hand-pinned row. The row is derived from what the variant carries. What was never read inside the image is Absent, sohub_cli_pathandvllm_versionbecame optional; the GLM canary's reader routes an absent CLI path to its existing unmodelled-runtime executor.deepseek_v41_measurement_profilemint. It fixes what V4.1 decides (fp8_ds_mla, and the file-backed Engram requestcpu_offloadtrue /dp_shared_memoryfalse). It takes what the capacity measurement owns as parameters.New
gunbc.spark.v41_group_a_launchis the one V4.1 call site:Names: distinct from every other arm (
gunbc-v41-tp4,gunbc-v41-tp4-{head,rank}.service), so a V4 pair or GLM unit still running reads as occupancy, never as this arm.Engram: the manifest file path and touch-receipt directory env (the overlay's
GUNBC_ENGRAM_ROW_STORE_MANIFEST/GUNBC_ENGRAM_TOUCH_RECEIPT_DIR). The row stores and manifest directory are mounted read-only at their own paths, and the receipt directory is writable. A per-host Engram staging script writes the manifest atomically before the apply's preflight, which requires mount sources to exist.Shape: one serve for the whole staircase (clever-gull-48):
max_model_len=v41_measurement_max_model_len, andmax_num_seqs= fold-max ofv41_staircase_concurrencies(16). It serves on the route's port (spark_pair_serve_port), because V4.1 succeeds V4 on the same serving subject.Admission (
v41_group_a_launch_admission, a pure fold;…_currentreads the authority log, never the source row). Each conjunct is its own located refusal:SuspendedForAuthorizedSuccessorwith a live lease (pair_serving_successor_may_launch).V41CandidateKeyed.candidate).runtime_artifactkey.admit_host_held_by_subject_among).Dry run (
v41_group_a_launch_dry_run) writestarget/v41-group-a-launch-dry-run.txt, containing:It exits 0 only when the plan rendered and admission holds.
Small shared moves:
std.measuregibibyte_ceiling(the Hub manifest's rounding now lives there, so one byte count cannot be rounded two ways), andgunbc.spark.model_snapshot_acquisitionspark_local_checkpoint_directory.Declared frontiers (DESIGN §3c): open PRs this consumes when they land
On main, each of these is a typed obligation that makes the dry run refuse by name, never a placeholder (proud-deer-538, Q2):
V41RuntimeProductionUnestablished; with no reading, the arm wall refuses "no registry reading exists"max_num_batched_tokens8192, utilization 8000 bp, prefix caching explicit offv41_measurement_launch_shape(clever-gull-48)v41_engram_store_manifestadmit_spark_host_unoccupiedSuspendedForAuthorizedSuccessor { exact_candidate_realization: PairRealizationKeyed { key: successor } }, lease observed at read timeEvidence
docker createargvs, every apply stage's script (preflight, preserve, apply, front door, commit, rollback, cleanup) and the compatibility subject key before and after the refactor: 57,535 bytes,cmpidentical. (Scratch, not enrolled: a digest literal copied from the tree would be a change detector, not an oracle. The enrolled GLM witnesses below cover the shape.)test.claim.spark.v41_group_a_launch_witness: 14/14 PASS (claim_batch). This includes the allowance pair: at 0.50 utilization the plan refuses and names the declared allowance, and without a profile the wall refuses. A zero-allowance mutation turns exactly the allowance claim red.--kv-cache-dtype fp8_ds_mla, both--engram-config.*flags, 262144 / 16, recipe parsers and tokenizer mode,--no-enable-prefix-caching, image = config digest. The Engram env and three mounts are present undergunbc-v41-tp4, with nogunbc-spark-pairorglm-nativeanywhere.plan,heldandstalewent red, and the positiveadmitcontrol stayed green.glm_group_b_launch_authority,serving_arm_kv_dtype,serving_fit_projectionandv41_checkpoint_materialize: 64 claims together with the 12 above.native_serving_roce_transportandserving_arm_launch: 28 claims.serving_runtime_capability_probe: all claims pass exceptthe_published_v41_image_admits_deepseek_v41_on_this_compute_capability, which fails identically on a clean origin/main worktree, so it is pre-existing.serving_arm_launch_witnessdid not resolve on main (every call lackeddeployment_env); it is repaired here and all its claims pass.🤖 Generated with Claude Code