Skip to content

Catalog memory components and the vLLM weight-offload split law (tier model, part 1 of 2) - #13331

Merged
gunbai-bot[bot] merged 3 commits into
mainfrom
session/bright-lark-561-catalog
Oct 5, 2026
Merged

gunbai-bot[bot] merged 3 commits into
mainfrom
session/bright-lark-561-catalog

Conversation

@gunbai-bot

@gunbai-bot gunbai-bot Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

What this is

Part 1 of 2 of the memory-tier model for the GH200 lane (node adhoc-e2818997-28d). It is split out of #13321 because the dashboard reviewer timed out three times on the combined diff (reviews 76027, 76033, 76057), each run ending at persona load. The parent asked for this split. #13321 keeps the arm_memory_fit migration and shrinks to that alone once this lands.

extdeps.systems.types: a system's memory components, addressable.

  • Before this change, catalog rows carried each memory's facts but gave no way to name "this memory of this system".
  • SystemMemoryComponent is the system's catalog name plus MemoryFacts plus GpuMemoryAddressability. Addressability takes one of three forms:
    • GpuAndCpuShareOnePool
    • GpuAddressesDirectly
    • GpuAddressesAcrossCoherentLink { per_direction }
  • Each row shape has its own projection, derived from fields that already exist (no new figures):
    • integrated_system_memory_components gives the GB10's one pool.
    • coherent_superchip_memory_components gives the GH200's HBM, plus its LPDDR5X across the C2C link at the row's own link rate.
  • A component's identity is the system name plus how the GPU reaches that memory. One system never reaches two memories the same way, so the pair names exactly one component.

extdeps.vllm.weight_offload: the engine's split law, cited at vLLM 8d09804c8.

  • vLLM's weight offloaders move whole parameters of decoder layers: make_layers calls wrap_modules.
  • UVA keeps no device copy.
  • Prefetch keeps prefetch_step device copies of each distinct offloaded parameter. The source is PrefetchOffloader.post_init, which builds StaticBufferPool(slot_capacity=prefetch_step).
  • The row states the condition the UVA law depends on: the MoE backend must not re-lay-out weights after loading. That holds for the TRITON / VLLM_CUTLASS arms of convert_to_fp8_moe_kernel_format. It does not hold for the DeepGEMM, FlashInfer, Marlin or AITER arms, which allocate new tensors on the device. The launch has to read this back; the row does not assume it.

Repairs after the side-chat REQUEST_CHANGES at 3303271

  1. Zero-copy UVA is an established realization, not a configured backend.
    • VllmWeightOffloadSplitLaw is now sole_constructor. Its only producer is vllm_weight_offload_split_standing.
    • For UVA, that producer reads three facts: is_uva_available(), VLLM_WEIGHT_OFFLOADING_DISABLE_UVA, and the realized VllmFp8MoeBackend. The backend is a closed coproduct of every Fp8MoeBackend arm, and each arm carries its convert_to_fp8_moe_kernel_format relayout fact.
    • It returns one of three standings:
      • an unread fact gives WeightOffloadSplitUnestablished, naming each obligation;
      • a relayout backend gives WeightOffloadSplitRefused;
      • UVA unavailable or disabled gives the functional_call fallback law, with one device copy;
      • only a full read with an in-place backend gives UvaZeroCopyView, with zero copies.
    • VllmPrefetchStep is sole_constructor behind vllm_prefetch_step, which refuses zero (upstream OffloadConfig.validate_offload_config).
    • Reds in test.claim.vllm_weight_offload_witness: UVA unavailable, UVA disabled, a relayout backend (DeepGEMM, Marlin), unread facts, prefetch step zero.
  2. Memory-component identity cannot alias.
    • SystemMemoryComponent is sole_constructor. Only the row-shape projections build it.
    • system_memory_component_same compares the whole fact: system, capacity, memory kind, and the complete addressability, including the coherent link's per-direction rate.
    • Reds in test.claim.system_memory_component_witness: a catalog row edited to the same system name with different GPU memory, and one with a different link rate. Neither aliases the original component.

Consumers (DESIGN §3c)

This PR is a declared frontier. Nothing in it is consumed in this diff. The named consumer is #13321, the second half of this split:

  • gunbc.spark.arm_memory_fit takes its memory tiers from these components. It joins them against the host's catalog memory, uses addressability to decide which tiers the engine envelope governs, and gives the Spark a one-tier subject through integrated_system_memory_components(nvidia_dgx_spark_4tb_catalog).
  • checkpoint_tier_split consumes VllmWeightOffloadSplitLaw.

The trigger is #13321 landing. If it does not land, both modules here are dangling and should be reverted.

🤖 Generated with Claude Code

… model, part 1 of 2)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
gunbc-ci-auto-heal and others added 2 commits October 5, 2026 05:46
…ry identity; reds for each

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…t hand tables

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gunbai-bot
gunbai-bot Bot added this pull request to the merge queue Oct 5, 2026
Merged via the queue into main with commit c82e497 Oct 5, 2026
4 checks passed
@gunbai-bot
gunbai-bot Bot deleted the session/bright-lark-561-catalog branch October 5, 2026 11:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants