Repository navigation
Catalog memory components and the vLLM weight-offload split law (tier model, part 1 of 2) - #13331
Merged
Merged
Conversation
… model, part 1 of 2) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ry identity; reds for each Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…t hand tables Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this is
Part 1 of 2 of the memory-tier model for the GH200 lane (node adhoc-e2818997-28d). It is split out of #13321 because the dashboard reviewer timed out three times on the combined diff (reviews 76027, 76033, 76057), each run ending at persona load. The parent asked for this split. #13321 keeps the
arm_memory_fitmigration and shrinks to that alone once this lands.extdeps.systems.types: a system's memory components, addressable.SystemMemoryComponentis the system's catalog name plusMemoryFactsplusGpuMemoryAddressability. Addressability takes one of three forms:GpuAndCpuShareOnePoolGpuAddressesDirectlyGpuAddressesAcrossCoherentLink { per_direction }integrated_system_memory_componentsgives the GB10's one pool.coherent_superchip_memory_componentsgives the GH200's HBM, plus its LPDDR5X across the C2C link at the row's own link rate.extdeps.vllm.weight_offload: the engine's split law, cited at vLLM 8d09804c8.make_layerscallswrap_modules.prefetch_stepdevice copies of each distinct offloaded parameter. The source isPrefetchOffloader.post_init, which buildsStaticBufferPool(slot_capacity=prefetch_step).convert_to_fp8_moe_kernel_format. It does not hold for the DeepGEMM, FlashInfer, Marlin or AITER arms, which allocate new tensors on the device. The launch has to read this back; the row does not assume it.Repairs after the side-chat REQUEST_CHANGES at 3303271
VllmWeightOffloadSplitLawis nowsole_constructor. Its only producer isvllm_weight_offload_split_standing.is_uva_available(),VLLM_WEIGHT_OFFLOADING_DISABLE_UVA, and the realizedVllmFp8MoeBackend. The backend is a closed coproduct of everyFp8MoeBackendarm, and each arm carries itsconvert_to_fp8_moe_kernel_formatrelayout fact.WeightOffloadSplitUnestablished, naming each obligation;WeightOffloadSplitRefused;functional_callfallback law, with one device copy;UvaZeroCopyView, with zero copies.VllmPrefetchStepissole_constructorbehindvllm_prefetch_step, which refuses zero (upstreamOffloadConfig.validate_offload_config).test.claim.vllm_weight_offload_witness: UVA unavailable, UVA disabled, a relayout backend (DeepGEMM, Marlin), unread facts, prefetch step zero.SystemMemoryComponentissole_constructor. Only the row-shape projections build it.system_memory_component_samecompares the whole fact: system, capacity, memory kind, and the complete addressability, including the coherent link's per-direction rate.test.claim.system_memory_component_witness: a catalog row edited to the same system name with different GPU memory, and one with a different link rate. Neither aliases the original component.Consumers (DESIGN §3c)
This PR is a declared frontier. Nothing in it is consumed in this diff. The named consumer is #13321, the second half of this split:
gunbc.spark.arm_memory_fittakes its memory tiers from these components. It joins them against the host's catalog memory, uses addressability to decide which tiers the engine envelope governs, and gives the Spark a one-tier subject throughintegrated_system_memory_components(nvidia_dgx_spark_4tb_catalog).checkpoint_tier_splitconsumesVllmWeightOffloadSplitLaw.The trigger is #13321 landing. If it does not land, both modules here are dangling and should be reverted.
🤖 Generated with Claude Code