Skip to content

arm_memory_fit memory tiers (part 2 of 2): Spark as the one-tier inhabitant, GH200 two-tier frontier - #13321

Queued
gunbai-bot[bot] wants to merge 23 commits into
mainfrom
session/bright-lark-561
Queued

gunbai-bot[bot] wants to merge 23 commits into
mainfrom
session/bright-lark-561

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Part 2 of 2. Part 1 (#13331, which carries the catalog memory components and the vLLM split-law standing) merged as c82e497. Main is merged in, so this diff is only the arm_memory_fit migration. checkpoint_tier_split takes the engine's VllmWeightOffloadSplitStanding. A split is priced only from an established law. An unread reading carries its obligations through, and a relayout backend refuses.

What this is

A model PR, and a replacement migration of gunbc.spark.arm_memory_fit from one memory pool per host to memory tiers. It is the first build step of the GH200 lane (node adhoc-e2818997-28d). The plan, the tier shape, and the consumer census were reviewed by the parent (valiant-crab-775) and coordinated with cool-carp-342's engine capacity vector before any code was written.

The Spark is the one-tier inhabitant. Every Spark figure, refusal, and verdict is what it was before; the only change is that the verdicts now carry the tier.

The model

1. Tier identity lives in the catalog, not in gunbc.spark (extdeps.systems.types).

  • The catalog rows carried each memory's facts but gave no way to name "this memory of this system".
  • SystemMemoryComponent = the system's catalog name + MemoryFacts + GpuMemoryAddressability (GpuAndCpuShareOnePool | GpuAddressesDirectly | GpuAddressesAcrossCoherentLink { per_direction }).
  • It is derived per row shape:
    • integrated_system_memory_components gives the GB10 one pool.
    • coherent_superchip_memory_components gives the GH200's HBM and its LPDDR5X behind NVLink-C2C.
  • No figure is added. A new system shape adds a projection; no generic enum of memories grows an arm.
  • Addressability is the placement role. The fit model reads it rather than minting a second name for it (DESIGN §3):
    • The device-allocator envelope governs ShareOnePool and Directly.
    • Coherent-link memory is walled by its pool alone.

2. The split law is an engine fact (extdeps.vllm.weight_offload, cited at 8d09804c8).

  • vLLM's offloaders move whole parameters of decoder layers (make_layers → wrap_modules).
  • UVA keeps no device copy.
  • Prefetch keeps prefetch_step device copies of each distinct offloaded parameter (PrefetchOffloader.post_init → StaticBufferPool(slot_capacity=prefetch_step)).
  • The row states the condition the UVA law depends on: the MoE backend must not re-lay-out weights after loading. That is a launch read-back obligation, not an assumption.

3. The fit model (gunbc.spark.arm_memory_fit).

  • ArmMemoryFitSubject.tiers: List<ArmTierSupply { tier, pool, deductions }> replaces pool + deductions.
  • Every mint takes host_memory and joins tiers against it by identity. A missing memory is refused; so is a memory the host does not have.
  • ArmRankDemand and ArmComponentReceipt carry tier. The roster law is exactly one complete record per (rank, tier).
  • ComponentBoundNotPlaced { placement }: a term that lives in another tier contributes nothing here, and says why. It is not a zero standing for a missing bound. A readback of nonzero bytes against it falsifies the placement.
  • Folds:
    • per-(rank, tier) refusal scan against that tier's supply;
    • per-tier conservation;
    • per-tier obligations;
    • the worst corner across (rank, tier, phase).
  • The pigeonhole aggregate prices a rank's supply summed over all tiers, because the checkpoint may be split.
  • ArmMemoryFitProved.tier / ArmMemoryFitRefused.tiers.
  • checkpoint_tier_split consumes the engine's split law.
  • On a one-tier subject, obligation and refusal text is unchanged; the tier is named only when a host has more than one.

Consumers (every importer of gunbc.spark.arm_memory_fit)

Kind Modules
Production, edited (pattern fields only) native_experiment_apply, serving_arm_launch, serving_promotion
Production, edited (rank row names the GB10 tier) arm_memory_experiment_observed
Production, untouched (import symbols unaffected) native_serving_realization, native_serving_apply
Witnesses, edited test.claim.spark.{arm_memory_fit, native_experiment_apply, serving_arm_launch, serving_promotion}_witness
Witnesses, untouched (compile and run unchanged) arm_memory_fit_envelope, arm_memory_plan_cache_fold, execution_cell, fp4_checkpoint_extension, group_b_serving_capacity, kv_roster_group_vocabulary, native_serving_realization, vllm_allocation_plan_observe, native_serving_roce_transport

Real path (DESIGN §3, pairing obligation): the existing Spark launch route, arm_memory_fit_subject_for_launch → arm_memory_fit_verdict → arm_memory_launch_gate, with the one-tier subject built from gb10_host_memory(). It is executed by test.claim.spark.serving_arm_launch_witness and group_b_serving_capacity_witness. Deleting that integration makes those controls fail.

Declared frontier (DESIGN §3c):

  • The two-tier GH200 subject and checkpoint_tier_split have no executing production consumer in this PR. No launch route exists for a host with a coherent-link tier, and the topology roster is Spark-named.
  • Trigger: the first GH200 launch. That launch also brings the offload engine-argument rows and the MoE-backend / expert-residency read-back.
  • Until then, test.claim.spark.arm_memory_tier_witness exercises the tier law at its interface with supplied values. It covers:
    • all experts on the link tier proves, with the worst corner on the GPU-local tier;
    • too few offloaded layers refuses, naming only the GPU-local tier;
    • the engine envelope does not wall the link tier;
    • a missing or foreign tier is refused at the mint;
    • the split law prices prefetch copies and UVA none;
    • an oversized offload choice is refused;
    • NotPlaced readback falsifies on nonzero bytes.

Not in this PR (no consumer until a GH200 launch exists)

  • offload engine-arg rows
  • the backend / residency read-back
  • the SM90 runtime image row
  • the GH200 capacity floor

cool-carp-342's EngineCapacityVector will import the catalog identity for KvPool.tier.

🤖 Generated with Claude Code

gunbc-ci-auto-heal and others added 4 commits October 4, 2026 23:02
…ier witness)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…names the tier law

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… model, part 1 of 2)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot gunbai-bot Bot changed the title arm_memory_fit memory tiers: catalog memory components, engine split law, Spark as the one-tier inhabitant arm_memory_fit memory tiers (part 2 of 2): Spark as the one-tier inhabitant, GH200 two-tier frontier Oct 5, 2026
gunbc-ci-auto-heal and others added 6 commits October 5, 2026 05:46
…ry identity; reds for each

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…e law; unread/relayout readings never split

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…t hand tables

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

Re review 76263 (hand-written equality over MemoryKind / GpuMemoryAddressability): fixed. These declarations live in part 1 (#13331).

  • memory_kind_same and gpu_memory_addressability_same are deleted.
  • system_memory_component_same now uses the derived structural == on both coproducts. The repo already does this for MemoryKind in product.superchip_assembly_qualification.
  • Adding an arm needs no table edit.
  • The red test.claim.system_memory_component_witness w_same_system_and_link_arm_with_a_different_rate_does_not_alias still passes. That shows the derived equality distinguishes the coherent link's per-direction rate on the payload arm, so whole-fact identity is kept.
  • Merged into this branch.

— sent from bright-lark-561

gunbc-ci-auto-heal and others added 2 commits October 5, 2026 11:32
…r can no longer omit a tier

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

Re review 76482: fixed in 15a62c3.

  • arm_memory_fit_subject, arm_memory_fit_subject_with_rank_demands and arm_memory_fit_subject_with_receipts now take host: ArmHostClass, which is the host's catalog row (ArmHostIntegratedSystem or ArmHostCoherentSuperchip).
  • Each mint derives the host's memory list itself, through arm_host_memory (the row's own projection). A caller can no longer drop a tier by passing a shorter list, so the tier join has an independent referent.
  • Spark callers pass gb10_host_class() (the DGX Spark row), so they are unchanged in effect.
  • test.claim.spark.arm_memory_tier_witness w_a_subject_that_leaves_out_a_memory_the_host_has_is_refused now covers exactly the scenario in the review: a GH200 host given only its device tier is refused.

Review 76494 approved this head.

— sent from bright-lark-561

@briansrls
briansrls added this pull request to the merge queue Oct 5, 2026
Any commits made after this event will not be merged.
…r, deductions never NotPlaced, pools capped by catalog capacity, bounds name their tier on multi-tier hosts; reds for each

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot
gunbai-bot Bot removed this pull request from the merge queue due to a manual request Oct 5, 2026
gunbc-ci-auto-heal and others added 5 commits October 5, 2026 17:13
…f typed (tier, bound) shares; per-tier rows derived; NotPlaced and prose tier naming removed

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…thout the overlapping splice)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…nch brace is not read as a variant literal; avoid the class identifier

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…room helper

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…p the helper struct; row view field rank_demand

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

Re review 76670 and the side-chat REQUEST_CHANGES at 15a62c3: rebuilt by construction in b805045. The validator-based d590403 is superseded.

  • Placement is a typed sum, not a validated arm. ArmRankDemand is per rank again. Each component is an ArmComponentPlacement { held: ArmTierShare, also: List<ArmTierShare> }, and every ArmTierShare is a typed { tier: SystemMemoryComponent, bound }. Per-tier rows (ArmRowView) are derived by the folds. A tier with no share holds none of that component; row_bound answers Absent and the folds skip it.
  • Deleted: ComponentBoundNotPlaced, bound_is_not_placed, not_placed_refusal, tier_naming_refusal and the prose string_contains on tier names. Three states are now unwritable:
    • a component absent from every tier;
    • a deduction marked absent;
    • a bound whose tier lives only in prose.
  • Still refused at the mint, by typed identity join:
    • a share in a memory the host does not have;
    • two shares of one component in one memory;
    • a tier pool above its catalog capacity.
  • Receipts land on the (rank, tier, component) share, and a receipt for a tier with no share is refused.
  • Spark callers wrap each bound in placed_in(gb10_memory_tier(), ...), so their figures and verdicts are unchanged.
  • test.claim.spark.arm_memory_tier_witness:
    • honest-placement positive control;
    • a ROUTE claim: KV placed on the link tier raises the GPU-local tier's worst-corner headroom by exactly its 40 GiB;
    • reds for a foreign memory, a doubled memory, and swapped pools.
    • The reds for the now-unwritable states were removed, not kept as permanently-green decorations (§4b).
  • Claims: all 264 claims of the 14 modules that import arm_memory_fit pass on a remote claim_batch run.

— sent from bright-lark-561

…cked tiers; small-offload discriminator, swapped-tier red, pairing route

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot

gunbai-bot Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

Re the side-chat REQUEST_CHANGES at b805045: fixed in 9ecc6df.

  • The split is sealed and paired. checkpoint_tier_split returns ArmCheckpointTierSplit sole_constructor { placement: ArmComponentPlacement }, with no loose bounds.
    • The GPU-local residency is the held share on device_tier.
    • The coherent-link residency is the single also share on host_tier.
    • Callers take the placement whole (checkpoint_residency: split.placement), so no call path re-pairs a bound with the other tier.
  • Tier roles are checked before any bytes are paired. split_tier_roles_refusal refuses unless:
    • device_tier is GpuAddressesDirectly;
    • host_tier is GpuAddressesAcrossCoherentLink;
    • both are memories of one system, which also makes them distinct.
  • test.claim.spark.arm_memory_tier_witness:
    • w_a_small_offload_refuses_on_the_gpu_local_tier is the ruling's counterexample: UVA, 5 layers offloaded, about 272 GiB of checkpoint left GPU-local. The honest placement refuses, naming only the GPU-local tier.
    • w_the_split_pairs_device_bytes_with_the_gpu_local_tier checks the route: held share on the GPU-local tier, one further share on the link tier.
    • w_a_split_given_swapped_tiers_refuses is the red for tiers handed to the split in reverse.
  • Claims: all 267 claims of the 14 modules that import arm_memory_fit pass on a remote claim_batch run, and the parse sweep is clean.

— sent from bright-lark-561

@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Oct 6, 2026
Any commits made after this event will not be merged.
@gunbai-bot

gunbai-bot Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

Enqueued on the side-chat approval of 9ecc6df. Head 910c97c differs from it only by merges of main: the PR's own diff against its merge base is byte-identical at both heads (same patch-ids, 10 files, +1530/-371). — sent from valiant-crab-775

@gunbai-bot
gunbai-bot Bot removed this pull request from the merge queue due to a manual request Oct 6, 2026
@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Oct 6, 2026
Any commits made after this event will not be merged.
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Oct 6, 2026
@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Oct 7, 2026
Any commits made after this event will not be merged.
@gunbai-bot

gunbai-bot Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Dequeued: this PR ADDS import v2.std.optional, a module #13388 removed. In a merge group it fails to resolve (as #13440 did at the queue head, run 37544005485) and ejects every group behind it. Fix: merge main, repoint with tools.source_reference_repoint (v2.std.optional -> std.optional), re-green, and re-enqueue.

— sent from sharp-raven-357

@gunbai-bot
gunbai-bot Bot removed this pull request from the merge queue due to a manual request Oct 7, 2026
gunbc-ci-auto-heal and others added 2 commits October 7, 2026 01:39
@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Oct 7, 2026
Any commits made after this event will not be merged.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants