Repository navigation
[unified-memory] Remove the kernel-page multiplier plumbing (5/7) - #40329
Merged
ch-wan merged 5 commits intoOct 2, 2026
Merged
Conversation
caihuali95
requested review from
BBuf,
DarkSharpness,
Fridge003,
HaiShaw,
HydraQYH,
Qiaolin-Yu,
Ying1123,
alphabetc1,
celve,
hanming-lu,
hebiao064,
hnyls2002,
huangtingwei9988,
hzh0425,
ispobock,
kpham-sgl,
merrymercy,
pyc96,
xiezhq-hermann,
yizhang2077 and
yuan-luo
as code owners
September 19, 2026 09:55
This was referenced Sep 19, 2026
Merged
This was referenced Sep 20, 2026
Merged
caihuali95
force-pushed
the
mainline/drop-kernel-page-multiplier
branch
from
September 24, 2026 02:35
c5d4058 to
5157236
Compare
caihuali95
requested review from
Edwardf0t1,
ch-wan and
fzyzcjy
as code owners
September 24, 2026 02:35
caihuali95
force-pushed
the
mainline/drop-kernel-page-multiplier
branch
2 times, most recently
from
September 25, 2026 05:49
db076c8 to
04df725
Compare
ch-wan
force-pushed
the
mainline/drop-kernel-page-multiplier
branch
from
September 28, 2026 19:53
04df725 to
5904873
Compare
ch-wan
requested review from
JustinTong0323,
sogalin,
wisclmy0611 and
zijiexia
as code owners
September 28, 2026 19:53
ch-wan
force-pushed
the
mainline/drop-kernel-page-multiplier
branch
4 times, most recently
from
October 1, 2026 06:21
156b037 to
2f5e7af
Compare
With token-major views the kernel-facing id is the physical token id, so the per-page block scale that the previous commit pinned to 1 has no remaining job. Remove it end to end; every site below multiplied by a constant 1: - MultiEndedAllocator: no `kernel_page_multiplier` kwarg or attribute, and `_translate_loc_fused` takes the page size as its stride. - The SWA and Mamba composites drop their multiplier properties. - KVIndexTranslator: no `_full/_swa_page_multiplier`; the read-table builders take no `multiplier=`; the SWA write loc derives with the page size alone. - kv_read_table: `entry = clamp(v2p[page], 0)`, the physical page. - create_mla_kv_page_table_for_dcp: no `mult` argument. Both translate implementations stay as they are. With the scale gone they compute the same id for any mapped virtual token, differing only in how they get there -- one Triton launch versus several torch ops -- and in what they do with a negative loc. Folding them is the next commit. The composite docstring that said tombstones map to -1 was wrong; both translates clamp them to 0, the padding sink. fa4, trtllm_mla, cutedsl_mla and tokenspeed_mla are admitted under --enable-unified-memory and supported by this change, but not yet validated: every validation run was on H20 (SM90), where their kernels do not run. The same holds for trtllm_mha's SM100 kernels (trtllm-gen context and decode) and for decode context parallelism (--dcp-size > 1). Reviewers with SM100-class hardware: please run them before merging.
…e-table call `create_mla_kv_page_table_for_dcp` no longer takes a page multiplier and `KVIndexTranslator` no longer has `full_page_multiplier`, but the aiter MLA backend's DCP verify path still passed it, so that path would fail on its first launch. Pass the same arguments as the trtllm_mla call site.
With the kernel-page multiplier gone, `write_loc_to_kernel_ids` still took a `stride` for the physical page, and its one caller passed the page size, so the kernel computed `phys * page_size + offset` under a name that promised a scale. Drop the argument; the tests that drove multipliers of 2, 3 and 7 now check the page-size form. The translator's module doc now names two id spaces, virtual and physical, and the comments that explained the removed scale state what the value is.
With the multiplier gone, the "single full-attention layer" tests (the multiplier-1 case) repeat the general ones: the MLA block-table test is a subset of test_block_table_matches_reference, and the fa3 one makes the same calls as the two translated-mapping tests. Fold their non-identity check into those tests and drop them, and stop citing the multiplier in the remaining docstrings.
TestPoolOwnership explained the guard with "ids up to num_pages * multiplier", a bound the multiplier's removal retired. The guard still holds for the reason that remains: the ids would be physical ids of a pool the runner does not own.
ch-wan
force-pushed
the
mainline/drop-kernel-page-multiplier
branch
from
October 2, 2026 00:11
2f5e7af to
21d1809
Compare
This was referenced Oct 2, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
5/7 — remove the kernel-page multiplier plumbing
Base: #40328 (4/7). 5 commits, 17 files, +117/−230.
What and why
3/7 pinned
kernel_page_multiplierto 1: under the token-major entry a kernel-facing idis the physical token id, so there is no per-page block scale left for an id to carry.
This PR deletes the plumbing that existed only to carry it.
Pure deletion. Nothing left behind computes differently — every site either multiplied
by a value that is now always 1, or passed it along.
What goes: the multiplier's keyword arguments and properties across the unified
allocators and the index translator, the
* multat every use site, the stride argumentsthat existed only to re-derive the scaled id, and the corresponding parameters in the
kernels that took them. The read-table build loses its scaling step and reduces to a
clamped table lookup. The aiter MLA backend's DCP verify path passed the multiplier to
create_mla_kv_page_table_for_dcptoo; it now passes the same arguments as trtllm_mla's.The write-loc translate kernel's
strideargument goes as well: its one caller passed thepage size, so the kernel computes
phys * page_size + offsetitself. The translator's moduledoc now describes two id spaces, virtual and physical.
Both translate implementations are left untouched here. With the multiplier gone,
translate_kv_locandtranslate_kv_loc_for_kernelcompute the same id for any in-rangenon-negative input, differing only in implementation (several torch ops vs one Triton
launch) and in how they treat a negative loc. Folding them is 6/7 — kept separate so
this PR carries no behaviour change at all.
Why it is not merged into 3/7
They are different kinds of change, and the split is deliberate:
files of deletion would dilute exactly the focus the split was made to create.
not find, reverting it alone leaves the layout intact — and the layout is what 4/7, 6/7,
7/7 and downstream work all stand on. Merged, that revert would take the layout with it.
Review shape
Net −113 lines across 17 files, eight of them tests whose edits follow the removal; test
cases that only exercised a multiplier above 1 are converted to the page-size form or
dropped, and the two single-full-attention-layer tests (the multiplier-1 case), which the
removal turned into duplicates, are folded into the general ones. Checking this PR is
mostly confirming that every deleted argument was provably 1.
Validation
Model families exercised — a hybrid of every pool shape the unified pool supports, plus a
non-hybrid control. Perf covers the gated-deltanet hybrid (Qwen3.5-9B), the MLA hybrid
(Kimi-Linear-48B-A3B), the sliding-window hybrid (gpt-oss-20b), the tri-pool hybrid
(Inkling and Inkling-Small), and dense Llama-3.1-8B as a control. Accuracy covers the same
set and additionally the Mamba hybrid, Falcon-H1-7B. Each model is run unified against a
baseline on the same model, backend and configuration, so every delta isolates the pool.
Median inter-token latency, unified vs baseline: +1.30% (measured run-to-run noise floor
for this metric: 1.9%). Median GSM8K delta at n=200: +0.00 pt (binomial delta-sigma ≈ 3.0 pt).
Re-validated on the previous base (0318a8d) after review feedback: the stack tip (7/7) gives +0.95% median inter-token latency over 40 paired configurations (median per-configuration shift against the previous run of the same 40: +0.12 pp) and −0.25 pt median GSM8K over 26, with the full unit scope green.
The median per-model perf shift against the preceding PR's run is −0.09 pp.
End to end
The admitted SM100 backends were run with unified memory at the stack tip; see #40331.
CI States
Latest PR Test (Base): ⏳ Run #36944701062
Latest PR Test (Extra): ⏳ Run #36944700814
Latest PR Test (AMD ROCm 10): ⏳ Run #36944701099