Skip to content

docs(memory): correct placement worked-example — gather is resident, 389 MB is lm_head - #1304

Merged
justinchuby merged 1 commit into
mainfrom
squad/1299-fix-placement-worked-example
Aug 18, 2026
Merged

justinchuby merged 1 commit into
mainfrom
squad/1299-fix-placement-worked-example

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Corrects the placement worked-example in docs/memory/MEMORY_MANAGEMENT_MODEL_DESIGN.md.

Why

The section used the embedding gather as its worked example of the placement lever ("moving 389,283,840 B to produce ~10 KB, ~78 ms/token to do no work"). The number is right, the tensor is not — anyone reading it reaches the same conclusion that greenlit a dead build. #1299 records the finding but an issue does not stop a reader of the doc; this fixes the passage itself.

What

Preserves the placement principle and the doc's measured/hardware/model honesty convention; corrects only the tensor identity and the now-known results.

Docs-only change. Relates to #1299, #1013, #1300, #994.

Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com

…389 MB is lm_head

The placement section used the embedding gather as its worked example of the
"move computation to the data" lever ("moving 389,283,840 B to produce ~10 KB,
~78 ms/token to do no work"). The number is right, the tensor is not.

Verified on qwen14b-zp (RTX 4060): `model.embed_tokens.qweight` is consumed by
`GatherBlockQuantized`, which is not a `LazyWeightBoundary`, so it is never paged
— it is resident and streams nothing per token (of 867 lazy-weight handles, zero
are named embed_tokens). The 389 MB/token observed streaming at key 919 (#945) is
`lm_head.weight` via `MatMulNBits` (152064 x 5120 x 0.5, INT4) — the vocab
projection, which does real arithmetic. Recorded in #1299.

Also folds in two now-settled results the section had marked open:
- lm_head GEMV placement is measured (#1013, x86): the CPU int4 kernel peaks at
  ~0.78 GB/s (scalar, no SIMD), so CPU lm_head is ~500 ms vs ~33 ms on GPU — the
  criterion inverts; host-placing lm_head is a net loss with the current kernel.
  Gap left open: mlas/accuracy_level=4 (prepacked) path, not measured by #1013.
- native CUDA-graph capture vs a per-token host excursion is checked (#1300):
  compatible only as an eager seam between captured segments (token-exact),
  illegal inside an active capture; seam price ~45–90 µs/token on the 4060. This
  is native-path only; #982 (plugin-EP interspersed-partition hang) is untouched.

Preserves the placement principle and the doc's measured/hardware/model
convention; corrects only the tensor identity and the now-known results.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby merged commit 0ebb868 into main Aug 18, 2026
7 of 16 checks passed
@justinchuby
justinchuby deleted the squad/1299-fix-placement-worked-example branch August 18, 2026 18:52
@codecov

codecov Bot commented Aug 19, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.88%. Comparing base (5949a3e) to head (a80c255).
⚠️ Report is 43 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff            @@
##           main    #1304       +/-   ##
=========================================
+ Coverage      0   80.88%   +80.88%     
=========================================
  Files         0      364      +364     
  Lines         0   160728   +160728     
  Branches      0   160728   +160728     
=========================================
+ Hits          0   130003   +130003     
- Misses        0    26070    +26070     
- Partials      0     4655     +4655     
Flag Coverage Δ
mlas 85.19% <ø> (?)
offline 80.80% <ø> (?)

Flags with carried forward coverage won't be shown. Click here to find out more.
see 364 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 gather/medium_f32_threads=1-internal/32768 4.37 µs 6.00 µs +37.2%
🔴 matmul/medium_generic_f16_threads=8/32x512x512 37.14 µs 50.52 µs +36.0%
🔴 add/medium_bf16_threads=1-internal/262144 96.72 µs 130.88 µs +35.3%
🔴 matmul/large_generic_f32_threads=8/32x1024x1024 5.22 ms 6.96 ms +33.3%
⚠️ grammar_masking/llguidance_compute_mask/32 82.62 µs 107.01 µs +29.5%
⚠️ matmul/medium_generic_f32_threads=8/32x512x512 1.53 ms 1.97 ms +28.9%
⚠️ matmul/small_generic_f16_threads=1/1x256x256 32.61 µs 41.84 µs +28.3%
⚠️ matmul/large_generic_f16_threads=8/32x1024x1024 101.72 µs 126.32 µs +24.2%
⚠️ gather/large_f16_threads=1-internal/131072 13.81 µs 16.94 µs +22.7%
⚠️ matmul/large_generic_bf16_threads=8/32x1024x1024 1.28 ms 1.56 ms +21.8%
⚠️ logit_processing/seven_processor_chain_per_step 339.85 µs 406.36 µs +19.6%
⚠️ matmul/large_generic_f16_threads=1/32x1024x1024 84.59 µs 99.49 µs +17.6%
⚠️ add/large_f32_threads=1-internal/4194304 767.47 µs 900.68 µs +17.4%
⚠️ matmul/medium_generic_f16_threads=1/32x512x512 40.15 µs 46.55 µs +15.9%
⚠️ sampling_latency/top_k_per_token 48.10 µs 55.70 µs +15.8%
⚠️ add/small_f32_threads=1-internal/1024 184.8 ns 213.8 ns +15.7%
⚠️ block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 61.28 µs 70.53 µs +15.1%
✅ matmul/small_generic_f32_threads=1/1x256x256 39.70 µs 45.59 µs +14.9%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 615.12 µs 699.68 µs +13.7%
✅ matmul/small_generic_bf16_threads=8/1x256x256 38.97 µs 43.53 µs +11.7%
✅ sampling_latency/min_p_per_token 193.78 µs 216.19 µs +11.6%
✅ gather/medium_bf16_threads=1-internal/32768 2.47 µs 2.76 µs +11.4%
✅ gather/medium_f16_threads=1-internal/32768 2.49 µs 2.78 µs +11.4%
✅ matmul/small_generic_bf16_threads=1/1x256x256 36.63 µs 40.21 µs +9.8%
✅ gather/small_f32_threads=1-internal/4096 665.7 ns 727.6 ns +9.3%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 92.47 µs 101.02 µs +9.2%
✅ gather/large_f32_threads=1-internal/131072 41.68 µs 45.34 µs +8.8%
✅ gather/large_bf16_threads=1-internal/131072 15.32 µs 16.41 µs +7.1%
✅ sampling_latency/greedy_per_token 3.03 µs 3.24 µs +6.9%
✅ reduce_mean/small_f32_threads=1-internal/4096 16.21 µs 17.22 µs +6.2%
✅ matmul/small_generic_f16_threads=8/1x256x256 37.65 µs 39.94 µs +6.1%
✅ add/medium_f32_threads=1-internal/262144 23.38 µs 24.73 µs +5.8%
✅ add/small_f16_threads=1-internal/1024 418.8 ns 440.9 ns +5.3%
✅ qwen3_sampling_processors/top_k_full_sort_baseline 1.95 ms 2.00 ms +2.7%
✅ qwen3_sampling_processors/top_k_partial_selection 130.22 µs 133.18 µs +2.3%
✅ tokenization/decode_tokens_per_second 6.02 ms 6.15 ms +2.1%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.09 ms 2.13 ms +1.9%
✅ sampling_latency/top_p_per_token 356.86 µs 360.79 µs +1.1%
✅ block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 72.00 µs 72.78 µs +1.1%
✅ kv_cache/alloc_dealloc_pages 38.84 µs 39.20 µs +0.9%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.07 ms 1.07 ms +0.1%
✅ add/small_bf16_threads=1-internal/1024 428.3 ns 422.5 ns -1.4%
✅ qwen3_sampling_processors/top_k_top_p_fast 701.72 µs 684.58 µs -2.4%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.85 ms 2.77 ms -2.6%
✅ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 384.82 µs 374.59 µs -2.7%
✅ add/large_f16_threads=1-internal/4194304 2.17 ms 2.11 ms -3.0%
✅ add/large_bf16_threads=1-internal/4194304 1.90 ms 1.84 ms -3.0%
✅ matmul/small_generic_f32_threads=8/1x256x256 57.96 µs 56.07 µs -3.3%
✅ gather/small_bf16_threads=1-internal/4096 521.7 ns 494.9 ns -5.1%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 684.76 µs 632.09 µs -7.7%
✅ tokenization/encode_tokens_per_second 425.01 µs 390.45 µs -8.1%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 557.85 µs 502.15 µs -10.0%
✅ gather/small_f16_threads=1-internal/4096 558.8 ns 499.4 ns -10.6%
✅ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 3.66 ms 3.25 ms -11.2%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 6.15 ms 5.43 ms -11.6%
✅ reduce_mean/medium_f32_threads=1-internal/65536 302.32 µs 263.91 µs -12.7%
✅ add/medium_f16_threads=1-internal/262144 126.60 µs 108.75 µs -14.1%
🟢 block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 980.16 µs 744.34 µs -24.1%
🟢 matmul/large_generic_f32_threads=1/32x1024x1024 12.51 ms 9.33 ms -25.4%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 4.32 3.72 4.68 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants