Skip to content

V4.1 TP4 launch on Group A (srv5-8): a modeled dry run through the mp arm launcher - #12378

Merged
gunbai-bot[bot] merged 5 commits into
mainfrom
session/zesty-ant-800
Sep 27, 2026
Merged

gunbai-bot[bot] merged 5 commits into
mainfrom
session/zesty-ant-800

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

DS4.1 critical path (work item adhoc-894b491d-00e): a modeled launch of the V4.1 four-rank (TP4) arm on Group A (srv5–8), as a dry run. Nothing in this PR launches on a real host. It has no wet entry; the only live read is the authority log. The design was reported to proud-deer-538 and approved before building, including one ruled change of route (below).

The route: Group B's mp-backend arm launcher, not the Ray pair

The brief asked to re-derive from the V4 Ray pair realization. That cannot run V4.1: the V4.1 image is upstream vLLM d2d649e6's vllm-openai target, and neither that revision's requirements/*.txt nor its docker/Dockerfile installs Ray. Ruled by proud-deer-538 (option A): use the launcher that already fits, a group-parameterized, Ray-free multi-host launcher (--distributed-executor-backend mp, --nnodes/--node-rank/--master-addr over the fabric rail, headless followers, FourSparkTp4):

  • gunbc.spark.serving_arm_launch plan_arm_launch_for plans it.
  • gunbc.spark.native_serving_realization renders the units.
  • gunbc.spark.native_serving_apply is the transactional apply (preflight / preserve / head-first apply / per-rank readback / commit or rollback).

Ray is not added to the V4.1 image, because that would change the image key. There is still one realization, parameterized (DESIGN §3); V4.1 is a new call site, like gunbc.spark.glm_group_b_launch.

(The brief also said the pair is TP2 over two hosts with its head on srv7. On main the pair is already TP = |Group A| = 4, with its head spark_pair_serving_head(FabricGroupA).)

What changes

The launcher stops being a GLM launcher. Each GLM literal it spelled moves to the authority that owns it:

  • Container name, rendezvous port, shm size, extra mounts become ArmContainerIdentity, a value the caller names. glm_native_arm_container carries Group B's unchanged values.
  • Tool/reasoning parsers move to the checkpoint row (CheckpointVllmParsers). Which parser decodes a model's output is a fact about the model. An optional tokenizer_mode renders no flag when absent.
  • Unit names, log paths and description become ArmUnitIdentity. glm_native_arm_units carries the GLM ones.
  • In native_serving_apply, each step carries its own log path and front door, and "head" is node_rank == 0 instead of a comparison against GLM's head-unit name. spark_native_arm_steps_of plans any arm. The stage scripts take the container name.

Upstream Engram request (extdeps.vllm.engine_args VllmEngramConfigRequest): vLLM d2d649e6's EngramConfig, rendered in upstream's own dotted form --engram-config.cpu_offload true --engram-config.dp_shared_memory false. vllm/utils/argparse_utils.py rewrites the dotted form into the nested JSON, so the words are plain and fit a systemd Exec line. On ServingLaunchProfile it enters the profile key only when explicit, so every existing GLM subject key is unchanged.

V4.1 rows in gunbc.spark.serving_arm:

  • DeepseekV41FlashDba1be0a checkpoint row. Placement is where v41_checkpoint_materialize installs the checkpoint; both now read one spark_local_checkpoint_directory. The parsers and tokenizer mode are the vLLM recipe's, which binds that recipe's "DSV41-1 launch argv binds …" frontier rows.
    • Its weight is the backbone (287 GiB); the Engram shards are served from row stores, not loaded as weights.
    • Engram page-cache allowance (proud-deer-538's correction): on a GB10 the page cache is the same unified pool, so the row declares resident_beyond_weights: PageCacheFromNonKvHeadroom. The fit wall charges it as the pool share the profile's --gpu-memory-utilization leaves outside the engine. That is the axis the measurement set for exactly this: at 0.80 it is 97 GiB across four ranks (about 24 GiB per rank). It is not zero, and it is not the whole table.
    • This is a declared bet, and v41_capacity_measurement v41_engram_resident_bets against v41_after_requests_resident confirms or falsifies it. Note: the stress model's bets (19/30/37/46 GiB per rank at x2/x4/x6/x16) exceed 24 GiB from x4 up, so the staircase's refault/PSI stop rules are what will discover whether that pressure matters.
    • Without a profile the charge is unknown, so serving_arm_admits refuses the row rather than fitting it on weights alone. The planner uses serving_arm_admits_under with its profile.
    • With the allowance: 384 GiB resident × 5/4 = 480 ≤ 484. It fits, by 4 GiB.
  • GunbcDeepseekV41EngramProduced { artifact, registry_reading } runtime variant. The V4.1 image is produced by the lane and admitted at run time, so its identity is a value, not a hand-pinned row. The row is derived from what the variant carries. What was never read inside the image is Absent, so hub_cli_path and vllm_version became optional; the GLM canary's reader routes an absent CLI path to its existing unmodelled-runtime executor.
  • deepseek_v41_measurement_profile mint. It fixes what V4.1 decides (fp8_ds_mla, and the file-backed Engram request cpu_offload true / dp_shared_memory false). It takes what the capacity measurement owns as parameters.

New gunbc.spark.v41_group_a_launch is the one V4.1 call site:

  • Names: distinct from every other arm (gunbc-v41-tp4, gunbc-v41-tp4-{head,rank}.service), so a V4 pair or GLM unit still running reads as occupancy, never as this arm.

  • Engram: the manifest file path and touch-receipt directory env (the overlay's GUNBC_ENGRAM_ROW_STORE_MANIFEST / GUNBC_ENGRAM_TOUCH_RECEIPT_DIR). The row stores and manifest directory are mounted read-only at their own paths, and the receipt directory is writable. A per-host Engram staging script writes the manifest atomically before the apply's preflight, which requires mount sources to exist.

  • Shape: one serve for the whole staircase (clever-gull-48): max_model_len = v41_measurement_max_model_len, and max_num_seqs = fold-max of v41_staircase_concurrencies (16). It serves on the route's port (spark_pair_serve_port), because V4.1 succeeds V4 on the same serving subject.

  • Admission (v41_group_a_launch_admission, a pure fold; …_current reads the authority log, never the source row). Each conjunct is its own located refusal:

    1. Group A is SuspendedForAuthorizedSuccessor with a live lease (pair_serving_successor_may_launch).
    2. The authorized successor key is exactly this V4.1 candidate (V41CandidateKeyed.candidate).
    3. The produced image was admitted against the candidate's runtime_artifact key.
    4. All four hosts are held by Group A's subject (admit_host_held_by_subject_among).
  • Dry run (v41_group_a_launch_dry_run) writes target/v41-group-a-launch-dry-run.txt, containing:

    • each host's create argv and unit text;
    • its Engram staging, preflight and apply scripts;
    • the admission verdict.

    It exits 0 only when the plan rendered and admission holds.

Small shared moves: std.measure gibibyte_ceiling (the Hub manifest's rounding now lives there, so one byte count cannot be rounded two ways), and gunbc.spark.model_snapshot_acquisition spark_local_checkpoint_directory.

Declared frontiers (DESIGN §3c): open PRs this consumes when they land

On main, each of these is a typed obligation that makes the dry run refuse by name, never a placeholder (proud-deer-538, Q2):

input owner today
produced V4.1 image (config digest) + a registry reading inside it build lane / #12353 V41RuntimeProductionUnestablished; with no reading, the arm wall refuses "no registry reading exists"
max_num_batched_tokens 8192, utilization 8000 bp, prefix caching explicit off #12371 v41_measurement_launch_shape (clever-gull-48) obligation naming #12371
Engram manifest text #12359 v41_engram_store_manifest obligation naming #12359
occupancy conjunct (a V4 unit or any GPU process is live, or memory is short, so refuse and never stop) #12351 admit_spark_host_unoccupied not yet a conjunct. TRIGGER: #12351 merges. SUFFICIENT FOR: a host whose V4 unit or any GPU process is live, or whose available memory is short, refuses this launch
host-effect claims the wet runner (the measurement's serve) the dry run lands no effect, so there is nothing to claim
the authority transition itself adhoc-f7cabf26-ca3 / #12370 (bright-ant-770) consumed as specified there: SuspendedForAuthorizedSuccessor { exact_candidate_realization: PairRealizationKeyed { key: successor } }, lease observed at read time

Evidence

  • Group B byte-identity (the regression control for the refactor). A scratch driver rendered all four Group B unit files, the four docker create argvs, every apply stage's script (preflight, preserve, apply, front door, commit, rollback, cleanup) and the compatibility subject key before and after the refactor: 57,535 bytes, cmp identical. (Scratch, not enrolled: a digest literal copied from the tree would be a change detector, not an oracle. The enrolled GLM witnesses below cover the shape.)
  • test.claim.spark.v41_group_a_launch_witness: 14/14 PASS (claim_batch). This includes the allowance pair: at 0.50 utilization the plan refuses and names the declared allowance, and without a profile the wall refuses. A zero-allowance mutation turns exactly the allowance claim red.
  • Mutation control: three independent defects were introduced at once (Engram request dropped from the mint, held-host conjunct removed, lease check removed). Exactly plan, held and stale went red, and the positive admit control stayed green.
  • Existing witnesses over the touched modules, all PASS:
    • glm_group_b_launch_authority, serving_arm_kv_dtype, serving_fit_projection and v41_checkpoint_materialize: 64 claims together with the 12 above.
    • native_serving_roce_transport and serving_arm_launch: 28 claims.
    • serving_runtime_capability_probe: all claims pass except the_published_v41_image_admits_deepseek_v41_on_this_compute_capability, which fails identically on a clean origin/main worktree, so it is pre-existing.
    • Note: serving_arm_launch_witness did not resolve on main (every call lacked deployment_env); it is repaired here and all its claims pass.

🤖 Generated with Claude Code

gunbc-ci-auto-heal and others added 5 commits September 26, 2026 22:11
…ich stops being GLM-only

The V4.1 image (upstream vLLM d2d649e6's vllm-openai target) ships no Ray, so it cannot run on the
V4 Ray pair. Per proud-deer-538's ruling, V4.1 is instead a new call site of the existing mp-backend
multi-host arm launcher (serving_arm_launch / native_serving_realization / native_serving_apply).
Changes:
- The launcher's GLM literals move to their owners: ArmContainerIdentity (container name,
  rendezvous port, shm size, extra mounts), CheckpointVllmParsers on the checkpoint row, and
  ArmUnitIdentity for unit names and logs.
- Group B renders byte-identical (57,535 bytes compared before and after).
- extdeps.vllm.engine_args VllmEngramConfigRequest adds vLLM's EngramConfig in its dotted argv form.
- serving_arm gains a V4.1 checkpoint row (backbone weight), a produced-runtime variant carrying
  its admitted identity, and deepseek_v41_measurement_profile.
- gunbc.spark.v41_group_a_launch plans all four hosts with distinct names, the Engram env and
  mounts, and per-host staging. It adds an admission fold that reads the authority log (successor
  may launch, candidate key, image key, held hosts), and a dry run that renders every host's
  create argv, unit and scripts.
- Inputs from open PRs (#12351, #12359, #12371, the produced image) are typed obligations that
  make the dry run refuse by name.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…llowance, not zero

On a GB10 the page cache is the same unified pool as the weights. So the V4.1 checkpoint row now
declares resident_beyond_weights = PageCacheFromNonKvHeadroom. The fit wall charges it as the pool
share the launch profile's memory fraction leaves outside the engine, which is the axis the capacity
measurement set for exactly this. It is not zero, and it is not the whole ~203 GB table.

The value is a bet: v41_engram_resident_bets, judged against v41_after_requests_resident, confirms or
falsifies it. Without a profile the charge is unknown, and serving_arm_admits refuses rather than
fitting on weights alone. The launch planner asks serving_arm_admits_under with its profile. GLM rows
declare NoResidentBeyondWeights, and Group B still renders byte-identical.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…t fns

Review 71644: the V4.1 Engram staging was a hand-assembled shell body. It is now three
gunbc.remote_operation PlannedRemoteSteps:
- the extdeps.tools.mkdir argv for the manifest and receipt directories;
- a gunbc.typed_remote_file_write sealed to the manifest directory.
The receipt directory needs no clearing, because the overlay writes each rank's receipt by temp
file and rename.

Required floor: test fns may not reference test code. Deleted:
- the aggregators in v41_group_a_launch_witness and glm_group_b_launch_authority_witness (refused
  by the floor);
- serving_arm_launch_witness's aggregator, together with its eight admitted test-reference debt
  rows in v1.compile (debt paid). The v1_compiler_compile.rs mirror is regenerated; it differs
  only by those eight rows.
Every claim still runs on its own.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Sep 27, 2026

Copy link
Copy Markdown
Contributor Author

Addressing review 71644 (the Engram staging was a hand-built shell script) in 7a35b1d, taking the reviewer's preferred fix rather than a marker: the staging is no longer a shell body. v41_engram_staging_steps is now three gunbc.remote_operation PlannedRemoteSteps:

  • the extdeps.tools.mkdir argv for the manifest directory;
  • the same for the touch-receipt directory;
  • a gunbc.typed_remote_file_write for the manifest, sealed to its own directory (FileTreeScope), so it lands as readback-verified bytes.

Two lines of the old script were dropped rather than modeled, because they were redundant:

  • clearing the receipt directory: the overlay's write_touch_receipt writes each rank<r>.json by temp file plus rename, so a receipt is always one whole document;
  • test -d on the row stores: the apply's preflight already checks every mount source.

The dry-run witness now asserts the manifest appears as a modeled write-file step and that no .next temp file from the old shell remains. The same commit also fixes the CI floor refusal: the floor refuses test fn aggregators that call other test fns, so they are deleted (each claim runs on its own), and the eight admitted test-reference debt rows for serving_arm_launch_witness's aggregator are retired in v1.compile with its regenerated mirror.

— sent from zesty-ant-800

@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Sep 27, 2026
Merged via the queue into main with commit b7eed9a Sep 27, 2026
5 checks passed
@gunbai-bot
gunbai-bot Bot deleted the session/zesty-ant-800 branch September 27, 2026 07:05
gunbai-bot Bot pushed a commit that referenced this pull request Sep 27, 2026
… MQ-1: a let value reads through body_lower_value_read beside main's carried annotation; the retired chunks stay retired and main's emitted_add chunk is kept; ledger receipts kept
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants