perf: unify metadata publication and reuse page tables and Triton kernels - #2399
Merged
Merged
Conversation
Publish persistent metadata at its first consumer boundary using checked direct copies or one packed H2D followed by scatter. Keep source lifetime, padding, stream identity, PP slots and explicit republication under one owner. Reject producer reentry before host writes and track actual packed-arena use. Share versioned page-table preparation across attention backends, including V4.1. Reuse unchanged maps, copy appended or reordered rows selectively, and record GPU revisions only after successful publication. Preserve worker append lineage without retaining page-id tuples or changing scheduler bookkeeping. Keep uniform sampling filters on the CPU, reuse independent TBO buffers, and combine Engram tentative cursors with snapshots. Group H2D regressions by contract and consumer; retain CPU page-map state tests separately.
Label staging reuse, preparation, model execution and postprocessing while the runner's torch profiler is active. Include batch token counts on forward markers and avoid record_function scopes when profiling is disabled.
Read the per-batch write width from the launch grid instead of specializing on each prefill length. Keep head_dim constexpr for the fixed layout and retain source bounds, ring indexing and padding behavior. Cover changing widths and odd head dimensions with GPU ring-write regressions.
Move runtime bounds and strides out of exact-value specialization across Gemma RMSNorm, MRoPE, cached KV quantization, MLA conversion, Qwen4 metadata, M3 context partitioning, and V4.1 packed gather. Read redundant counts from the launch grid while retaining model geometry and useful bounded variants. Preserve vectorized KV accesses with fixed row widths and int64 lengths. Unroll only the existing small-input reduction configuration and avoid allocating unused single-pass scratch. Keep all bounds and padding guards. Add one focused suite covering output correctness, tail boundaries, large integer signatures, and compiled-kernel reuse. Relevant GPU regressions: 176 passed. Black, Ruff, and diff checks passed for this change.
Contributor
🏷️ CI GuideRuns automatically on every eligible PR before approval:
Heavy model tests:
|
yhl-amd
pushed a commit
that referenced
this pull request
Sep 28, 2026
…ether Every backend kept `block_regions` and `block_tensor_views` as two lists paired by hand (DSV4 even collected a third list of sources to zip later). `KVTransferTensors.add_block_region(tensor, semantic_role=...)` now appends both from one tensor: the region's addresses and a zero-copy `uint8 [num_units, 1, unit_bytes]` alias of exactly those bytes, cut from a larger allocation when `total_bytes` says so. DSV4, MHA, the MHA draft and MLA all publish through it. MHA and the draft already published this shape. MLA's views change from `[n, block_size, width]` to the same byte form; LMCache stores the same bytes in the same order either way, and neither its object key nor ATOM's model namespace depends on view shape, so existing cache entries stay valid. The MLA builder test's module stubs had drifted since #2399 (`atom.utils.block_tables`); they are fixed so the test runs again. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
valarLip
pushed a commit
that referenced
this pull request
Sep 28, 2026
* feat(offload): support native DSv4 checkpoints with LMCache MP Save and restore DSv4 PAGE KV together with the matching recurrent STATE checkpoint through a standalone LMCache multiprocess server. The lmcache_mp connector selects its implementation from backend capabilities: backends that publish a PagedStateCheckpointSpec and execute_paged_state_copies use the native PAGE/STATE path, others keep the PAGE-only transport. Saves lease immutable READY checkpoint PAGE units instead of snapshotting the live SLOT; restores reserve fresh units and adopt them as a local READY checkpoint. Also: - allow single-host DP and DP-attention, scoping MP sessions per replica; - release source PAGE blocks per chunk and reacquire a finished request's still-resident prefix at save admission; - allow partial release of recurrent-state requests only when the connector guarantees an independent state lease; - read offload env knobs through atom.utils.envs; - refuse PAGE-copied state backends that lack the native contract. Requires LMCache dev@05fc77a (LMCache/LMCache#5132), pinned on main by #2395. Start the server with --null-block-id -1 --separate-object-groups. Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(offload): pass --eviction-policy to the LMCache MP server example The pinned LMCache (05fc77a) requires --eviction-policy on `lmcache server`; without it the documented command exits 2 during argument parsing. Reported-by: kvnloo Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(lmcache_mp): skip native saves shorter than OFFLOAD_MIN_SAVE_TOKENS Native MP stored every READY checkpoint boundary, including whole short prompts, although a prefix below OFFLOAD_MIN_LOAD_TOKENS (8192 by default, equal to OFFLOAD_MIN_SAVE_TOKENS) can never be loaded back. On DSv4-Pro TP8 with a 1K/1K workload at concurrency 64, offload stored 257 such prefixes and cost 12.4% output throughput; skipping them stores none and leaves 1.6-2.2%, within run-to-run noise. 16K prompts still save as before. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(offload): keep early release safe without a BlockManager; one save threshold Review follow-ups (#2250 §1, §4). - Unbound schedulers (the vLLM plugin, where vLLM owns the blocks) cannot reacquire a finished request's prefix, so teardown again freezes the block table and leases the unemitted suffix, and the final save reads exactly those leased blocks. The late-save reacquire path now runs only when a BlockManager is bound, instead of keying off an empty block table that the plugin never clears. - The shared base no longer reads OFFLOAD_MIN_SAVE_TOKENS, so dense and hybrid behave as before this PR. A late save with nothing savable resident now retires the request instead of emitting an empty save forever. - Native MP applies OFFLOAD_MIN_SAVE_TOKENS to the absolute boundary for normal and late saves alike, so a long request's short tail is still stored, and the floor can never reach boundary 0. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(lmcache_mp): bound unprovable transfers; fence native restores Review follow-ups (#2250 §3, §7). - A submission that raises, a restore that raises, or a restore event whose query raises may still be touching engine memory, so its lease is kept -- but no longer forever. _UncertainSubmission now fails the transfer after lmcache.mp.uncertain_transfer_timeout_s (default: twice lmcache.mp.mq_timeout) with a warning, so the save or load settles and its lease, budget, pending slot and descriptor slot are released. Before, one ZMQ hiccup could stop all native saves and loads and keep EngineCore busy-looping. - PAGE-only MP treats a raising save submission the same way instead of as a definite failure. The failure reported no quiescence, so under #2339 the save was never retried and its lease was never released. The now-dead _immediate_save_failures set is removed. - Native restores are fenced both ways: the restore stream waits for work already on the compute stream (write-after-read on the SLOT's previous occupant), and every step's compute stream waits on in-flight restore events (read-after-write, including SLOT relocations in build()). The fence is on the GPU, so the host does not block. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(lmcache_mp): end sessions after the last save; unwind failed admissions Review follow-ups (#2250 §7). - The MP scheduler ended a request's session in request_finished, but with early release its final save is emitted and submitted under that session afterwards. Session ends are now deferred until no save of the request is tracked or in flight. They are swept on retirement and at every metadata build. - Native save admission returns the checkpoint lease and budget charge if anything after the acquire raises; the pin is never timeout-reclaimable, so it used to leak for the process lifetime. Load admission computes the boundary hash before reserving units, and releases the units, the budget and any suspended local restore if the rest raises. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(lmcache_mp): keep draft PAGE regions out of the native image Review follow-ups (#2250 §5, §7). - DSv4 now publishes paged_state_region_count, the number of its own PAGE regions. A draft with a pool of its own appends regions after them, and the native layout validates checkpoint coverage and builds STATE aliases from the leading regions only. Draft rows are registered as ordinary PAGE KV and never folded into the checkpoint image. Before, DSv4 with such a draft failed registration with "PAGE regions do not cover the native PAGE unit". - Auto rank collapse is off when a DSpark draft's backend owns a KV pool: that draft publishes per-rank PAGE regions, so the complete PAGE object is not replicated and registration would otherwise reject the requested collapse. - The native layout registers its PAGE group as uint8 views, like the PAGE-only path and like its own STATE aliases. LMCache's ROCm raw-pointer fallback cannot express FP8 through the CUDA array interface. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(attention): take the descriptor slot in DSv4.1 paged-state copies Review follow-up (#2250 §5). descriptor_slot reached the base class, DSv4 and GDN but not DSv4.1, whose execute_paged_state_copies(stores, restores) would raise TypeError inside the native restore's exception handler and be reported as a failed restore. StateCopies now keeps an independent pinned staging buffer and upload fence per descriptor slot, like the other backends. A new AST sweep test fails if any attention backend's implementation lacks the parameter. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(offload): seed native and late-save hash chains from cache_seed Review follow-up (#2250 §2). The native boundary-hash fallback and BlockManager.acquire_offload_prefix both started their chains at -1, while every BlockManager chain starts at seq.cache_seed, which multimodal requests set. For those requests every native save was skipped, every restored checkpoint was published under a key no loader looks up, and late saves found nothing resident. The native fallback now extends its chain through the new BlockManager.prefix_hash_chain (the manager's own _chain_to), so both _boundary_hash branches share one seed, algorithm and token slicing. acquire_offload_prefix seeds from seq.cache_seed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(offload): fail closed on disabled paged-state checkpoints and composite state release Review follow-ups (#2250 §6). - The MP scheduler shell picked the native-state scheduler whenever the checkpoint coordinator existed, but the coordinator exists without prefix caching and then can never hold a READY image, so offload silently did nothing. A PAGE-only fallback would restore KV under stale recurrent state. Binding now fails with a message naming --enable-prefix-caching. - MultiConnectorScheduler.can_partially_deallocate_state now requires every still-deferring sub to guarantee state safety. Before, one declaring sub was enough, and the state slot could be recycled while a non-declaring sub (e.g. a dense PAGE sub) still needed the request alive. The test now separates all-of from any-of. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * perf(lmcache_mp): cache the native save frontier; reserve restore slots early Review follow-ups (#2250 §5, §8). - Native _save_frontier walked the prompt checkpoint by checkpoint for every tracked request on every scheduler step. PageUnitCheckpointStore now has a generation that is bumped whenever the READY set can change, and the answer is cached per request on (frontier, floor, generation). - Restore descriptor slots >= 1 were first allocated as pinned memory on the connector thread mid-serving. The native worker now reserves them at registration through the new builder hook reserve_checkpoint_descriptors (base builders allocate their descriptor buffers, DSv4.1 its per-slot staging). - Document why single-host DP replicas share one (model_name, worker_id, world_size) identity. It is the content-addressed storage namespace, while the server registers GPU memory per instance_id, refcounts layout descriptors per (model_name, world_size), and _mp_session_id scopes sessions per replica. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(offload): document every offload env var; guard it with a test Review follow-ups (#2250 §8, §9). - docs/environment_variables.md listed only OFFLOAD_MAX_PENDING_SAVES and the LMCache pin timeout, and still said they bypass atom.utils.envs. All 16 ATOM-owned offload knobs are now documented with type, default, precedence and invalid-value behavior. A new test fails if an OFFLOAD_*/LMCACHE_* var registered in envs.py is missing from the reference. - README and envs.py describe OFFLOAD_MIN_SAVE_TOKENS as the native absolute boundary for normal and late saves, and the README documents the uncertain-transfer bound. - Test that a restore raising after it took a descriptor slot fails the load and returns the slot once the uncertainty bound expires. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(lmcache_mp): make the raising-restore test independent of CUDA On a CPU-only torch, torch.cuda.Event is a dummy class that raises when it is built, which happens before the restore takes its descriptor slot, so the test's "slot is held until the bound" assertion failed in CI. Stub the Event so the restore fails deterministically at the stream fence, after the slot is taken, on every runner. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(lmcache_mp): release transfer memory only on a report, fail stop at a deadline The MP path answered "what settles a transfer whose outcome is unprovable" three ways: never (native checkpoint pins, a zeroed abandon timeout), after 600 s (the uncertainty bound, which then freed load destinations and PAGE sources a server might still be using), and never-unbounded (a future whose poll keeps raising). One rule now: memory under an MP transfer is released only on a terminal report, and a transfer still not terminal after `lmcache.mp.transfer_deadline_s` (default 1200 s) raises `LMCacheTransferUnprovable`, stopping the engine. - Worker: every pending PAGE-only and native load/save carries its start time and is checked against the deadline, which bounds a raising submission, a future whose poll keeps raising, and a restore whose event cannot be queried. A descriptor that cannot be built is provably unsent and still fails immediately; only the transport call is unprovable. - Scheduler: a watchdog over dispatched saves, loads and native checkpoint sources catches a completion that never arrives (lost report, silent TP rank), with a margin so the worker fails first. - `save_abandon_timeout_s` is inherited again, so the engine-wide stalled save, state-pin and orphan-load-slot reclaimers stay on; MP opts out through its own `abandon_save` / `reclaim_stale_leases` instead. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(offload): lease the final save source when prefix caching is off The late final save reacquired a finished request's prefix through the hash index whenever a BlockManager was bound. With prefix caching off nothing is indexed, so the lookup found nothing and the save was dropped without a warning -- the combination the offload README ships for `lmcache_offload`. Hash reacquire now requires a bound manager with prefix caching; otherwise teardown leases the unemitted suffix of the frozen table, as on main. A late save that finds part of its prefix evicted is counted in `truncated_late_saves`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(offload): reacquire a final save source only after a partial release `protected_block_ids` marked the request finished before the scheduler decided whether it could release it partially. When per-request state made that unsafe, the request was deferred whole and kept its table, but the next metadata build still took the late-save path and claimed every block a second time through the hash index (dropping the save if a hash was missing). `activate_block_leases`, which only the partial-release branch calls, now marks the release, and only that mark selects the reacquire path. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(lmcache_mp): keep the native STATE image pinned until the STORE terminal The worker released a save's checkpoint image as soon as any PAGE chunk milestone reached the boundary. Chunk milestones report PAGE token ranges; nothing in LMCache promises that a chunk completes only after every engine group registered at it, STATE groups included, has been read. The pinned LMCache emits no chunk events, so this path never ran, but it would have been unsound on the first server that did. The STATE source-safe channel is removed: the image pin is released at the STORE terminal, and chunk milestones release PAGE leases only. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(envs): read an empty offload knob as its default everywhere `OFFLOAD_PUBLICATION_TIMEOUT_S=` made `float('')` raise out of connector init, and `OFFLOAD_COPY_WORKERS` / `OFFLOAD_LOAD_WORKERS` died with a bare `invalid literal for int()`. The offload section now states one policy: unset or empty is the default; a set but unusable value warns and falls back for knobs that only tune reuse, and is rejected at startup, naming the variable, for widths, sizes and timeouts. The transfer-mode comment no longer lists `engine_driven` as valid. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(lmcache_mp): keep failed adoptions releasable and surface broken invariants - `adopt_transfer_units` popped the transfer record before adopting; if adoption raised, the units were reserved but unreachable by any release path. The record is now removed only after adoption returns. - A native restore that loses the publish race to an identical image was dropped silently; it is now logged as a deduplication. - The worker's own pending-save refusal is unreachable while the scheduler enforces the same bound; if it ever fires it now logs an error instead of passing as an ordinary save failure. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(lmcache_mp): run a native restore end to end and fix drifted copy doubles The restore-copy doubles were positional-only, so the production call with `descriptor_slot=` would have raised inside `_begin_restore` and been reported as a failed restore; the success tests replaced `_begin_restore` outright, so no test ran the happy path. The doubles now take the production signature, and a new test drives a terminal retrieve through `get_finished` into the real `_begin_restore`, then asserts the copy's descriptor slot, that the load finishes only once the restore event is done, and that the slot is returned. The CUDA stand-ins are shared with the fence test. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * perf(lmcache_mp): invalidate a request's save frontier only on its own checkpoints The per-request `_save_frontier` memo was keyed on the store-wide generation, so any request's publish or eviction forced every tracked request to rescan its prompt on the next step. The store now keeps a bounded log of which prefix hash each generation bump touched; a request rescans only when a change hits one of the boundaries it scanned (its answer or anything above it), or when the log cannot tell. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(attentions): one fenced descriptor staging for every checkpoint copier DSv4.1 kept its own `_DescriptorStaging`, a third copy of the per-slot pinned descriptor pool the base builder shares with DSv4 and GDN, and the only one that fenced reuse: a non-blocking H2D reads its pinned rows when the stream reaches it, so refilling a buffer whose last upload is still queued rewrites the descriptor under an earlier copy. `DescriptorStaging` in `pool_layout/paged_state_copy.py` now owns the per-slot buffers and the fence, and the base builder (DSv4, GDN) and DSv4.1's `StateCopies` both use it, so the base path gains the fence too. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(lmcache_mp): validate PAGE views once for both registrations `build_native_state_mp_layout` and `_build_cache_views` each validated the backend-published PAGE views, and had already drifted: PAGE-only required full contiguity and compared total bytes against the view, native checked inner contiguity plus the block stride against the region. Both now call `validate_page_views` (new `mp/page_views.py`), which applies the stricter union -- shape, tight block-major stride, unit and total bytes, aliasing, forward indexing, one device -- and, for native, that a block's physical slots divide the block size. Each builder keeps only what is its own: PAGE-only rejects stateful fields; native checks the image coverage. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(lmcache_mp): split the native worker's completion poll `get_finished` was 91 lines at six levels of nesting, with three near-identical recovery blocks. It is now a loop over `_poll_native_save` and `_poll_native_load`; the load's retrieve-then-restore progression lives in `_advance_native_load`, and both unprovable-restore paths share `_hold_unprovable_restore`. No behaviour change. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(lmcache_mp): split mp/backend.py by responsibility `mp/backend.py` had grown to 1332 lines holding configuration, PAGE view validation, transfer bookkeeping, lookups and both connector halves. It is split without behaviour change into: - `deployment.py`: config validation, TP/DP topology and rank collapse, model namespace, server adapters - `transfer.py`: operation identity, terminal detection, the transfer deadline and `LMCacheTransferUnprovable` - `lookup.py`: the scheduler's lookup client and read-lock bookkeeping - `page_views.py`: PAGE-only `_build_cache_views`, beside the shared view validation it uses - `worker.py` / `scheduler.py`: the PAGE-only connector halves The native modules, the public shell and the tests import from the new homes; tests patch each name where it is looked up. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(offload): keep clock reclaim off MP requests; warn on truncated late saves - With the abandon window inherited again, a slow but legitimate MP save deferred past it was "abandoned" (a no-op for MP) and, a minute later, logged as a wedged P/D send -- both long before MP's own transfer deadline. A connector can now answer `waits_for_transfer_report(seq)`; the stalled-save reclaim skips such requests. LMCache MP answers True, the in-process shell forwards (default False), and the composite answers True only if every sub still deferring the request does, so P/D sends and in-process saves keep their abandon path. - A late save that finds part of its prefix evicted now logs a warning with the running count, not a debug line: early release returns the tail's blocks at teardown, and this is the cost of that trade. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(kv_transfer): publish each PAGE region and its byte view together Every backend kept `block_regions` and `block_tensor_views` as two lists paired by hand (DSV4 even collected a third list of sources to zip later). `KVTransferTensors.add_block_region(tensor, semantic_role=...)` now appends both from one tensor: the region's addresses and a zero-copy `uint8 [num_units, 1, unit_bytes]` alias of exactly those bytes, cut from a larger allocation when `total_bytes` says so. DSV4, MHA, the MHA draft and MLA all publish through it. MHA and the draft already published this shape. MLA's views change from `[n, block_size, width]` to the same byte form; LMCache stores the same bytes in the same order either way, and neither its object key nor ATOM's model namespace depends on view shape, so existing cache entries stay valid. The MLA builder test's module stubs had drifted since #2399 (`atom.utils.block_tables`); they are fixed so the test runs again. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(kv_transfer): make a PAGE unit one value, its region and view together `add_block_region` paired a region and its view at one call site, but the pair was only a convention: both lists stayed public and mutable, the constructor still took them separately, DSV4 and MLA used a throwaway `KVTransferTensors` as a scratch builder, the draft merge extended the two lists by hand, and tensor code lived in the torch-free `types` contract. - `PageRegion(region, view)` is a frozen value; `KVTransferTensors.pages` holds them. `block_regions` and `block_tensor_views` are read-only tuples derived from it, so no code can add to one without the other. - `page_region(tensor, ...)` in the new `disaggregation/page_region.py` builds one from the owning tensor (zero-copy byte view, contiguity and size checks); `types.py` no longer touches torch. - Builders pass `pages=[...]` straight to the final object: no scratch instance. The address-only producer (Qwen4 exp) publishes `PageRegion(region)` with no view, which LMCache MP refuses as before. - `merge_pages(other)` replaces the draft merge in `ModelRunner`, carrying the gcd replication rule with it. `validate_page_views` stays on the LMCache MP side: it is the consumer checking what it is handed, not a second copy of construction. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Honglie Yi <hyi@crsuse2-m2m-v2-020.us-east2-a.compute.internal> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
jiayyu
added a commit
that referenced
this pull request
Sep 29, 2026
Squash of the 33 commits of fpz/mixed_mla_dispatch_v4 (history preserved in fpz/mixed_mla_dispatch_v4_backup_20260929), rebased onto main e59d422. Ported onto main's checked H2D publication (#2399): - The mixed prefill bank is now a bound publication slot with its own owner, groups and event (ModelRunner.build_h2d_slot), one per forward_vars ring slot. Swapping forward_vars alone made the prefill producers publish the runner's (stale) sources through h2d_groups. - cu_seqlens_q has one writer per buffer per epoch: publish_cu_seqlens_q publishes the decode segment's spans on a mixed step, and mixed_prefill_bank_active the prefill segment's into the bank. prepare_mixed no longer re-uploads it. - Dense decode block tables go through block_table_state; V4 merges positions on the device instead of re-staging a published buffer. - Mixed deferred input ids use fill_deferred_decode_ids over the decode region. - The TBO prefill-segment rebuild swaps owner and groups with the bank. - _flash_attn_prefill takes `out` (FlyDSL FP8 copies into it); the chunk helpers take the decode-first prefill budget; _settle_prefill_chunks ends a chunk at num_tokens, per main. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Forward preparation uploads many small metadata buffers and repacks page tables even when the page mapping is unchanged. Changing batch sizes and prefill lengths also creates new Triton specializations for values that only control runtime bounds.
This PR publishes metadata at its first consumer under a shared ownership contract, reuses page-table state across attention backends, and reuses Triton kernels across dynamic lengths.
ATOM_H2D_BACKEND=directremains the default;packedcombines eligible metadata into one H2D followed by scatter.Changes
Add checked publication with source lifetime, stream/device identity, semantic padding, and failure handling owned by the existing runner slots. A shared producer decorator checks write permission before host data can be overwritten. Packed arena ownership follows actual packed transfers, including mixed direct/packed publication.
Share versioned page-table preparation across common attention, MHA/MLA, V4, and V4.1. Unchanged mappings avoid page-ID scans and uploads; appends and request reordering update the affected rows. CPU preparation and successful GPU publication have separate revisions. PP/TBO buffers retain independent state, and V4.1 uses the common entry point.
Integrate runner, attention, draft/GDN, and Engram consumers at their first-use boundaries. Keep uniform sampling filters on the CPU, reuse independent TBO buffers, and preserve graph padding and fixed destination storage.
Remove exact-length specialization from SWA writes and 13 additional Triton kernels. Preserve model geometry, tiles, and useful bounded alignment variants. Cached KV quantization keeps vectorized accesses and supports inputs exceeding 2^31 elements; only the existing small-input reduction configuration uses two-way loop unrolling.
Add
ATOM::runner stage annotations when profiling is active, document the publication contract, and organize H2D regressions by owner/consumer responsibility.Bind and retain the model-runner TCPStore before spawning workers. All managed ranks connect as clients, removing the gap between probing a free port and binding it later. Single-node runs allocate the port with
port=0; multi-node runs share an explicit fixed port. RapidServe retains separate prefill/decode stores, and standalone ModelRunner callers keep their existing rendezvous path.Validation
Validation used an AMD Instinct MI355X with PyTorch 2.9.1+ROCm7.1.1 and Triton 3.8.0.
Optionalannotation warnings and four exception-handling warnings in the manager.The suites overlap and their counts should not be added. Performance checks use 1M context and operator/metadata microbenchmarks, not end-to-end model throughput. In a controlled shared-buffer comparison, 4096-token KV quantization measured 11.95 → 11.02 μs after the reduction adjustment (11.14 μs before the Triton changes). Some small dynamic shapes retain overhead, including about 0.30 μs for 257-token KV; the accepted shared-page-table path also retains about 2.4 μs in its B256 single-request subsequent-chunk case. Full-model serving, quality, and throughput validation was not rerun.