Repository navigation
chore(squad): record native-path placement capture-compat finding - #1300
Merged
Merged
Conversation
…em 3) Decision drop: per-token device excursion is compatible with native CUDA graph capture as an eager seam (~47-90 us/token, RTX 4060); the doc's embedding-gather placement target is resident on qwen14b-zp (streamed 389MB/token is lm_head). Scribe merges into the shared decisions log. See #1299. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
added a commit
that referenced
this pull request
Aug 18, 2026
…389 MB is lm_head (#1304) Corrects the placement worked-example in `docs/memory/MEMORY_MANAGEMENT_MODEL_DESIGN.md`. ## Why The section used the embedding gather as its worked example of the placement lever ("moving 389,283,840 B to produce ~10 KB, ~78 ms/token to do no work"). The number is right, the tensor is not — anyone reading it reaches the same conclusion that greenlit a dead build. #1299 records the finding but an issue does not stop a reader of the doc; this fixes the passage itself. ## What - **Tensor identity corrected.** Verified on `qwen14b-zp` (RTX 4060): `model.embed_tokens.qweight` is consumed by `GatherBlockQuantized`, which is **not** a `LazyWeightBoundary`, so it is never paged — it is **resident** and streams nothing per token (of 867 lazy-weight handles, zero are named `embed_tokens`). The 389,283,840 B (`152064 x 5120 x 0.5`, INT4) streaming at key 919 (#945) is `lm_head.weight` via `MatMulNBits` — the vocab projection, which does real work. The gather remains a correct *statement* of the `F ~ 0` principle; it is simply already resident, so there is nothing to move. - **lm_head GEMV result folded in (#1013, x86, info-only reference):** the CPU int4 `MatMulNBits` kernel peaks at ~0.78 GB/s (scalar, no SIMD), so CPU `lm_head` is ~500 ms vs ~33 ms on GPU — the criterion **inverts**; host-placing `lm_head` is a net loss with the current kernel. Remaining gap: the `mlas`/`accuracy_level=4` prepacked path #1013 did not measure. - **Native CUDA-graph capture interaction resolved (#1300):** a per-token device→host→device excursion is compatible with native graph capture **only as an eager seam between captured segments** (token-exact) and **illegal inside** an active capture; seam price ~45–90 µs/token on the 4060. Native path only — **#982 (plugin-EP interspersed-partition hang) is untouched and must not be read as cleared.** Preserves the placement principle and the doc's measured/hardware/model honesty convention; corrects only the tensor identity and the now-known results. Docs-only change. Relates to #1299, #1013, #1300, #994. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #1300 +/- ##
=========================================
+ Coverage 0 80.70% +80.70%
=========================================
Files 0 364 +364
Lines 0 160728 +160728
Branches 0 160728 +160728
=========================================
+ Hits 0 129710 +129710
- Misses 0 26361 +26361
- Partials 0 4657 +4657
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
|
| Status | Scenario | Base | PR | Change |
|---|---|---|---|---|
sampling_latency/top_p_per_token |
439.01 µs | 531.19 µs | +21.0% | |
qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline |
4.12 ms | 4.98 ms | +20.9% | |
grammar_masking/llguidance_compute_mask/32 |
94.32 µs | 111.90 µs | +18.6% | |
kv_cache/alloc_dealloc_pages |
44.66 µs | 52.85 µs | +18.3% | |
qwen3_sampling_processors/top_k_full_sort_baseline |
2.44 ms | 2.88 ms | +17.9% | |
sampling_latency/min_p_per_token |
241.90 µs | 279.67 µs | +15.6% | |
| ✅ | gather/large_f16_threads=1-internal/131072 |
16.99 µs | 19.41 µs | +14.2% |
| ✅ | logit_processing/seven_processor_chain_per_step |
372.57 µs | 420.23 µs | +12.8% |
| ✅ | qwen3_sampling_processors/top_k_top_p_fast |
751.00 µs | 845.42 µs | +12.6% |
| ✅ | sampling_latency/top_k_per_token |
63.49 µs | 70.68 µs | +11.3% |
| ✅ | qwen3_sampling_processors/top_k_top_p_full_sort_baseline |
7.28 ms | 8.09 ms | +11.1% |
| ✅ | qwen3_sampling_processors/top_k_partial_selection |
162.56 µs | 180.49 µs | +11.0% |
| ✅ | gather/medium_f32_threads=1-internal/32768 |
4.53 µs | 4.91 µs | +8.4% |
| ✅ | matmul/medium_generic_f16_threads=8/32x512x512 |
47.35 µs | 51.24 µs | +8.2% |
| ✅ | qwen3_sampling_processors/top_p_fast_after_top_k |
586.57 µs | 628.44 µs | +7.1% |
| ✅ | reduce_mean/small_f32_threads=1-internal/4096 |
16.80 µs | 17.78 µs | +5.8% |
| ✅ | add/large_bf16_threads=1-internal/4194304 |
1.86 ms | 1.96 ms | +5.4% |
| ✅ | matmul/small_generic_f16_threads=8/1x256x256 |
40.08 µs | 42.16 µs | +5.2% |
| ✅ | add/small_f16_threads=1-internal/1024 |
519.3 ns | 543.0 ns | +4.6% |
| ✅ | gather/large_bf16_threads=1-internal/131072 |
18.47 µs | 19.19 µs | +3.9% |
| ✅ | add/medium_f16_threads=1-internal/262144 |
118.61 µs | 121.55 µs | +2.5% |
| ✅ | matmul/medium_generic_bf16_threads=8/32x512x512 |
660.49 µs | 675.94 µs | +2.3% |
| ✅ | matmul/large_generic_f32_threads=8/32x1024x1024 |
7.09 ms | 7.24 ms | +2.2% |
| ✅ | matmul/large_generic_bf16_threads=1/32x1024x1024 |
2.40 ms | 2.45 ms | +2.1% |
| ✅ | gather/large_f32_threads=1-internal/131072 |
52.28 µs | 53.25 µs | +1.8% |
| ✅ | matmul/medium_generic_f32_threads=1/32x512x512 |
2.73 ms | 2.78 ms | +1.8% |
| ✅ | block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 |
139.74 µs | 142.19 µs | +1.8% |
| ✅ | block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 |
1.37 ms | 1.39 ms | +1.5% |
| ✅ | reduce_mean/medium_f32_threads=1-internal/65536 |
277.86 µs | 281.89 µs | +1.5% |
| ✅ | matmul/large_generic_f32_threads=1/32x1024x1024 |
11.04 ms | 11.19 ms | +1.4% |
| ✅ | gather/medium_f16_threads=1-internal/32768 |
2.81 µs | 2.84 µs | +1.2% |
| ✅ | gather/small_f32_threads=1-internal/4096 |
757.6 ns | 765.6 ns | +1.1% |
| ✅ | block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 |
588.09 µs | 593.02 µs | +0.8% |
| ✅ | tokenization/decode_tokens_per_second |
7.54 ms | 7.60 ms | +0.8% |
| ✅ | matmul/small_generic_f16_threads=1/1x256x256 |
39.11 µs | 39.40 µs | +0.7% |
| ✅ | matmul/large_generic_f16_threads=8/32x1024x1024 |
120.87 µs | 121.73 µs | +0.7% |
| ✅ | tokenization/encode_tokens_per_second |
445.81 µs | 447.93 µs | +0.5% |
| ✅ | matmul/small_generic_bf16_threads=8/1x256x256 |
41.73 µs | 41.85 µs | +0.3% |
| ✅ | matmul/medium_generic_f32_threads=8/32x512x512 |
1.85 ms | 1.85 ms | +0.2% |
| ✅ | gather/small_bf16_threads=1-internal/4096 |
535.2 ns | 534.3 ns | -0.2% |
| ✅ | matmul/small_generic_f32_threads=1/1x256x256 |
48.20 µs | 48.08 µs | -0.3% |
| ✅ | matmul/medium_generic_f16_threads=1/32x512x512 |
40.40 µs | 40.21 µs | -0.5% |
| ✅ | matmul/large_generic_f16_threads=1/32x1024x1024 |
97.05 µs | 96.54 µs | -0.5% |
| ✅ | add/large_f16_threads=1-internal/4194304 |
1.89 ms | 1.88 ms | -0.6% |
| ✅ | matmul/large_generic_bf16_threads=8/32x1024x1024 |
2.40 ms | 2.37 ms | -1.1% |
| ✅ | add/small_f32_threads=1-internal/1024 |
239.5 ns | 236.5 ns | -1.3% |
| ✅ | gather/small_f16_threads=1-internal/4096 |
557.6 ns | 549.2 ns | -1.5% |
| ✅ | block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 |
120.54 µs | 118.69 µs | -1.5% |
| ✅ | add/large_f32_threads=1-internal/4194304 |
971.95 µs | 955.44 µs | -1.7% |
| ✅ | matmul/medium_generic_bf16_threads=1/32x512x512 |
660.79 µs | 647.35 µs | -2.0% |
| ✅ | matmul/small_generic_bf16_threads=1/1x256x256 |
40.32 µs | 39.45 µs | -2.2% |
| ✅ | block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 |
121.74 µs | 118.50 µs | -2.7% |
| ✅ | gather/medium_bf16_threads=1-internal/32768 |
2.94 µs | 2.86 µs | -2.7% |
| ✅ | sampling_latency/greedy_per_token |
3.72 µs | 3.59 µs | -3.6% |
| ✅ | matmul/small_generic_f32_threads=8/1x256x256 |
59.04 µs | 56.27 µs | -4.7% |
| ✅ | reduce_mean/large_f32_threads=1-internal/262144 |
1.19 ms | 1.12 ms | -5.8% |
| ✅ | add/medium_bf16_threads=1-internal/262144 |
128.21 µs | 116.91 µs | -8.8% |
| 🟢 | add/small_bf16_threads=1-internal/1024 |
632.8 ns | 512.5 ns | -19.0% |
| 🟢 | add/medium_f32_threads=1-internal/262144 |
36.39 µs | 28.63 µs | -21.3% |
Visual flags:
Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 5.28 4.61 5.70 }
What this cannot catch
- Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
- Sub-threshold regressions that compound over multiple PRs
- Performance changes that only manifest under GPU execution
- Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Decision inbox drop for the Scribe recording item-3 findings. See #1299. Native path only; #982 untouched.