Skip to content

[unified-memory] Remove the kernel-page multiplier plumbing (5/7) - #40329

Merged
ch-wan merged 5 commits into
sgl-project:mainfrom
caihuali95:mainline/drop-kernel-page-multiplier
Oct 2, 2026
Merged

ch-wan merged 5 commits into
sgl-project:mainfrom
caihuali95:mainline/drop-kernel-page-multiplier

Conversation

@caihuali95

@caihuali95 caihuali95 commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

This is one of seven PRs split out of #38592. That PR was +2105/−1366 across 41
files, twelve of them attention backends on the serving hot path, and it interleaved
hardening, a layout change and dead-code deletion in a single diff — a reviewer could not
tell which hunk was safe-by-construction and which changed behaviour. #38592 is retained as the token-major
layout PR (3/7), which is the core of the series.

Stack — each PR is based on the one above it, and every one is green on its own:

# PR what it does size
1 #40326 derive KV row addresses from strides, not shapes 21 files, +733/−118
2 #40327 build paged KV views through one helper 13 files, +704/−47
3 #38592 token-major layout 36 files, +1215/−1108
4 #40328 mark the write loc physical and check it at every write door 45 files, +935/−132
5 #40329 remove the kernel-page multiplier plumbing 17 files, +117/−230
6 #40330 stride the KV translate kernel and route every translate through it 11 files, +361/−159
7 #40331 name the fused KV translate for what it computes 28 files, +119/−121

Review order is the table order. 1 and 2 are no-ops on today's contiguous pools (2 also
fixes how HND pools with more than one KV head are read); 3 is where the behaviour changes
and deserves the most attention; 4 adds the physical write-loc check; 5 is deletion; 6
consolidates the translate; 7 is naming only.

5/7 — remove the kernel-page multiplier plumbing

Base: #40328 (4/7). 5 commits, 17 files, +117/−230.

What and why

3/7 pinned kernel_page_multiplier to 1: under the token-major entry a kernel-facing id
is the physical token id, so there is no per-page block scale left for an id to carry.
This PR deletes the plumbing that existed only to carry it.

Pure deletion. Nothing left behind computes differently — every site either multiplied
by a value that is now always 1, or passed it along.

What goes: the multiplier's keyword arguments and properties across the unified
allocators and the index translator, the * mult at every use site, the stride arguments
that existed only to re-derive the scaled id, and the corresponding parameters in the
kernels that took them. The read-table build loses its scaling step and reduces to a
clamped table lookup. The aiter MLA backend's DCP verify path passed the multiplier to
create_mla_kv_page_table_for_dcp too; it now passes the same arguments as trtllm_mla's.
The write-loc translate kernel's stride argument goes as well: its one caller passed the
page size, so the kernel computes phys * page_size + offset itself. The translator's module
doc now describes two id spaces, virtual and physical.

Both translate implementations are left untouched here. With the multiplier gone,
translate_kv_loc and translate_kv_loc_for_kernel compute the same id for any in-range
non-negative input, differing only in implementation (several torch ops vs one Triton
launch) and in how they treat a negative loc. Folding them is 6/7 — kept separate so
this PR carries no behaviour change at all.

Why it is not merged into 3/7

They are different kinds of change, and the split is deliberate:

  • 3/7 is the feature, and its middle commit is where a reviewer must think hard. Seventeen
    files of deletion would dilute exactly the focus the split was made to create.
  • They are independently revertible. If this removal turns out to break a consumer we did
    not find, reverting it alone leaves the layout intact — and the layout is what 4/7, 6/7,
    7/7 and downstream work all stand on. Merged, that revert would take the layout with it.

Review shape

Net −113 lines across 17 files, eight of them tests whose edits follow the removal; test
cases that only exercised a multiplier above 1 are converted to the page-size form or
dropped, and the two single-full-attention-layer tests (the multiplier-1 case), which the
removal turned into duplicates, are folded into the general ones. Checking this PR is
mostly confirming that every deleted argument was provably 1.

Validation

Model families exercised — a hybrid of every pool shape the unified pool supports, plus a
non-hybrid control. Perf covers the gated-deltanet hybrid (Qwen3.5-9B), the MLA hybrid
(Kimi-Linear-48B-A3B), the sliding-window hybrid (gpt-oss-20b), the tri-pool hybrid
(Inkling and Inkling-Small), and dense Llama-3.1-8B as a control. Accuracy covers the same
set and additionally the Mamba hybrid, Falcon-H1-7B. Each model is run unified against a
baseline on the same model, backend and configuration, so every delta isolates the pool.

Median inter-token latency, unified vs baseline: +1.30% (measured run-to-run noise floor
for this metric: 1.9%). Median GSM8K delta at n=200: +0.00 pt (binomial delta-sigma ≈ 3.0 pt).

Re-validated on the previous base (0318a8d) after review feedback: the stack tip (7/7) gives +0.95% median inter-token latency over 40 paired configurations (median per-configuration shift against the previous run of the same 40: +0.12 pp) and −0.25 pt median GSM8K over 26, with the full unit scope green.

The median per-model perf shift against the preceding PR's run is −0.09 pp.

End to end

The admitted SM100 backends were run with unified memory at the stack tip; see #40331.


CI States

Latest PR Test (Base): ⏳ Run #36944701062
Latest PR Test (Extra): ⏳ Run #36944700814
Latest PR Test (AMD ROCm 10): ⏳ Run #36944701099

@caihuali95
caihuali95 force-pushed the mainline/drop-kernel-page-multiplier branch from c5d4058 to 5157236 Compare September 24, 2026 02:35
@github-actions github-actions Bot added the amd label Sep 24, 2026
@caihuali95
caihuali95 force-pushed the mainline/drop-kernel-page-multiplier branch 2 times, most recently from db076c8 to 04df725 Compare September 25, 2026 05:49
@github-actions github-actions Bot added the hicache Hierarchical Caching for SGLang label Sep 25, 2026
@ch-wan
ch-wan force-pushed the mainline/drop-kernel-page-multiplier branch from 04df725 to 5904873 Compare September 28, 2026 19:53
@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Sep 28, 2026
@ch-wan
ch-wan force-pushed the mainline/drop-kernel-page-multiplier branch 4 times, most recently from 156b037 to 2f5e7af Compare October 1, 2026 06:21
Caihua Li and others added 5 commits October 1, 2026 17:10
With token-major views the kernel-facing id is the physical token id, so the
per-page block scale that the previous commit pinned to 1 has no remaining
job. Remove it end to end; every site below multiplied by a constant 1:

- MultiEndedAllocator: no `kernel_page_multiplier` kwarg or attribute, and
  `_translate_loc_fused` takes the page size as its stride.
- The SWA and Mamba composites drop their multiplier properties.
- KVIndexTranslator: no `_full/_swa_page_multiplier`; the read-table builders
  take no `multiplier=`; the SWA write loc derives with the page size alone.
- kv_read_table: `entry = clamp(v2p[page], 0)`, the physical page.
- create_mla_kv_page_table_for_dcp: no `mult` argument.

Both translate implementations stay as they are. With the scale gone they
compute the same id for any mapped virtual token, differing only in how they
get there -- one Triton launch versus several torch ops -- and in what they do
with a negative loc. Folding them is the next commit.

The composite docstring that said tombstones map to -1 was wrong; both
translates clamp them to 0, the padding sink.

fa4, trtllm_mla, cutedsl_mla and tokenspeed_mla are admitted under
--enable-unified-memory and supported by this change, but not yet validated:
every validation run was on H20 (SM90), where their kernels do not run. The
same holds for trtllm_mha's SM100 kernels (trtllm-gen context and decode)
and for decode context parallelism (--dcp-size > 1). Reviewers with
SM100-class hardware: please run them before merging.
…e-table call

`create_mla_kv_page_table_for_dcp` no longer takes a page multiplier and
`KVIndexTranslator` no longer has `full_page_multiplier`, but the aiter MLA
backend's DCP verify path still passed it, so that path would fail on its
first launch. Pass the same arguments as the trtllm_mla call site.
With the kernel-page multiplier gone, `write_loc_to_kernel_ids` still took a
`stride` for the physical page, and its one caller passed the page size, so
the kernel computed `phys * page_size + offset` under a name that promised a
scale. Drop the argument; the tests that drove multipliers of 2, 3 and 7 now
check the page-size form.

The translator's module doc now names two id spaces, virtual and physical,
and the comments that explained the removed scale state what the value is.
With the multiplier gone, the "single full-attention layer" tests (the
multiplier-1 case) repeat the general ones: the MLA block-table test is
a subset of test_block_table_matches_reference, and the fa3 one makes
the same calls as the two translated-mapping tests. Fold their
non-identity check into those tests and drop them, and stop citing the
multiplier in the remaining docstrings.
TestPoolOwnership explained the guard with "ids up to num_pages *
multiplier", a bound the multiplier's removal retired. The guard still
holds for the reason that remains: the ids would be physical ids of a
pool the runner does not own.
@ch-wan
ch-wan force-pushed the mainline/drop-kernel-page-multiplier branch from 2f5e7af to 21d1809 Compare October 2, 2026 00:11
@ch-wan
ch-wan merged commit 9855e9c into sgl-project:main Oct 2, 2026
24 of 30 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

amd blackwell SM100/SM120 documentation Improvements or additions to documentation hicache Hierarchical Caching for SGLang jit-kernel memory-pool

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants