Repository navigation
docs(memory): correct placement worked-example — gather is resident, 389 MB is lm_head - #1304
Merged
Merged
Conversation
…389 MB is lm_head
The placement section used the embedding gather as its worked example of the
"move computation to the data" lever ("moving 389,283,840 B to produce ~10 KB,
~78 ms/token to do no work"). The number is right, the tensor is not.
Verified on qwen14b-zp (RTX 4060): `model.embed_tokens.qweight` is consumed by
`GatherBlockQuantized`, which is not a `LazyWeightBoundary`, so it is never paged
— it is resident and streams nothing per token (of 867 lazy-weight handles, zero
are named embed_tokens). The 389 MB/token observed streaming at key 919 (#945) is
`lm_head.weight` via `MatMulNBits` (152064 x 5120 x 0.5, INT4) — the vocab
projection, which does real arithmetic. Recorded in #1299.
Also folds in two now-settled results the section had marked open:
- lm_head GEMV placement is measured (#1013, x86): the CPU int4 kernel peaks at
~0.78 GB/s (scalar, no SIMD), so CPU lm_head is ~500 ms vs ~33 ms on GPU — the
criterion inverts; host-placing lm_head is a net loss with the current kernel.
Gap left open: mlas/accuracy_level=4 (prepacked) path, not measured by #1013.
- native CUDA-graph capture vs a per-token host excursion is checked (#1300):
compatible only as an eager seam between captured segments (token-exact),
illegal inside an active capture; seam price ~45–90 µs/token on the 4060. This
is native-path only; #982 (plugin-EP interspersed-partition hang) is untouched.
Preserves the placement principle and the doc's measured/hardware/model
convention; corrects only the tensor identity and the now-known results.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1304 +/- ##
=========================================
+ Coverage 0 80.88% +80.88%
=========================================
Files 0 364 +364
Lines 0 160728 +160728
Branches 0 160728 +160728
=========================================
+ Hits 0 130003 +130003
- Misses 0 26070 +26070
- Partials 0 4655 +4655
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
🔴 Benchmark Regression DetectedComparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).
Visual flags: Host infoWhat this cannot catch
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Corrects the placement worked-example in
docs/memory/MEMORY_MANAGEMENT_MODEL_DESIGN.md.Why
The section used the embedding gather as its worked example of the placement lever ("moving 389,283,840 B to produce ~10 KB, ~78 ms/token to do no work"). The number is right, the tensor is not — anyone reading it reaches the same conclusion that greenlit a dead build. #1299 records the finding but an issue does not stop a reader of the doc; this fixes the passage itself.
What
qwen14b-zp(RTX 4060):model.embed_tokens.qweightis consumed byGatherBlockQuantized, which is not aLazyWeightBoundary, so it is never paged — it is resident and streams nothing per token (of 867 lazy-weight handles, zero are namedembed_tokens). The 389,283,840 B (152064 x 5120 x 0.5, INT4) streaming at key 919 (A pinned, never-re-admitted page of the int4 lm_head weight goes stale — #886 isolated to one tensor #945) islm_head.weightviaMatMulNBits— the vocab projection, which does real work. The gather remains a correct statement of theF ~ 0principle; it is simply already resident, so there is nothing to move.MatMulNBitskernel peaks at ~0.78 GB/s (scalar, no SIMD), so CPUlm_headis ~500 ms vs ~33 ms on GPU — the criterion inverts; host-placinglm_headis a net loss with the current kernel. Remaining gap: themlas/accuracy_level=4prepacked path Add roofline_gemv: measure the real CPU int4 lm_head GEMV against DRAM (#994) — re-measured, ~3x below roofline, not ~63x #1013 did not measure.Preserves the placement principle and the doc's measured/hardware/model honesty convention; corrects only the tensor identity and the now-known results.
Docs-only change. Relates to #1299, #1013, #1300, #994.
Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com