Skip to content

chore(squad): record native-path placement capture-compat finding - #1300

Merged
justinchuby merged 1 commit into
mainfrom
chore/placement-decision-drop
Aug 18, 2026
Merged

justinchuby merged 1 commit into
mainfrom
chore/placement-decision-drop

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Decision inbox drop for the Scribe recording item-3 findings. See #1299. Native path only; #982 untouched.

…em 3)

Decision drop: per-token device excursion is compatible with native CUDA graph
capture as an eager seam (~47-90 us/token, RTX 4060); the doc's embedding-gather
placement target is resident on qwen14b-zp (streamed 389MB/token is lm_head).
Scribe merges into the shared decisions log. See #1299.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby merged commit c88b3c8 into main Aug 18, 2026
7 of 16 checks passed
@justinchuby
justinchuby deleted the chore/placement-decision-drop branch August 18, 2026 18:46
justinchuby added a commit that referenced this pull request Aug 18, 2026
…389 MB is lm_head (#1304)

Corrects the placement worked-example in
`docs/memory/MEMORY_MANAGEMENT_MODEL_DESIGN.md`.

## Why
The section used the embedding gather as its worked example of the
placement lever ("moving 389,283,840 B to produce ~10 KB, ~78 ms/token
to do no work"). The number is right, the tensor is not — anyone reading
it reaches the same conclusion that greenlit a dead build. #1299 records
the finding but an issue does not stop a reader of the doc; this fixes
the passage itself.

## What
- **Tensor identity corrected.** Verified on `qwen14b-zp` (RTX 4060):
`model.embed_tokens.qweight` is consumed by `GatherBlockQuantized`,
which is **not** a `LazyWeightBoundary`, so it is never paged — it is
**resident** and streams nothing per token (of 867 lazy-weight handles,
zero are named `embed_tokens`). The 389,283,840 B (`152064 x 5120 x
0.5`, INT4) streaming at key 919 (#945) is `lm_head.weight` via
`MatMulNBits` — the vocab projection, which does real work. The gather
remains a correct *statement* of the `F ~ 0` principle; it is simply
already resident, so there is nothing to move.
- **lm_head GEMV result folded in (#1013, x86, info-only reference):**
the CPU int4 `MatMulNBits` kernel peaks at ~0.78 GB/s (scalar, no SIMD),
so CPU `lm_head` is ~500 ms vs ~33 ms on GPU — the criterion
**inverts**; host-placing `lm_head` is a net loss with the current
kernel. Remaining gap: the `mlas`/`accuracy_level=4` prepacked path
#1013 did not measure.
- **Native CUDA-graph capture interaction resolved (#1300):** a
per-token device→host→device excursion is compatible with native graph
capture **only as an eager seam between captured segments**
(token-exact) and **illegal inside** an active capture; seam price
~45–90 µs/token on the 4060. Native path only — **#982 (plugin-EP
interspersed-partition hang) is untouched and must not be read as
cleared.**

Preserves the placement principle and the doc's measured/hardware/model
honesty convention; corrects only the tensor identity and the now-known
results.

Docs-only change. Relates to #1299, #1013, #1300, #994.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Co-authored-by: justinchuby <223556219+Copilot@users.noreply.github.com>
@codecov

codecov Bot commented Aug 19, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.70%. Comparing base (2abbbdb) to head (8b3ad44).
⚠️ Report is 46 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff            @@
##           main    #1300       +/-   ##
=========================================
+ Coverage      0   80.70%   +80.70%     
=========================================
  Files         0      364      +364     
  Lines         0   160728   +160728     
  Branches      0   160728   +160728     
=========================================
+ Hits          0   129710   +129710     
- Misses        0    26361    +26361     
- Partials      0     4657     +4657     
Flag Coverage Δ
mlas 85.22% <ø> (?)
offline 80.61% <ø> (?)

Flags with carried forward coverage won't be shown. Click here to find out more.
see 364 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown

⚠️ Benchmark Change Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
⚠️ sampling_latency/top_p_per_token 439.01 µs 531.19 µs +21.0%
⚠️ qwen3_sampling_processors/top_p_full_sort_after_top_k_baseline 4.12 ms 4.98 ms +20.9%
⚠️ grammar_masking/llguidance_compute_mask/32 94.32 µs 111.90 µs +18.6%
⚠️ kv_cache/alloc_dealloc_pages 44.66 µs 52.85 µs +18.3%
⚠️ qwen3_sampling_processors/top_k_full_sort_baseline 2.44 ms 2.88 ms +17.9%
⚠️ sampling_latency/min_p_per_token 241.90 µs 279.67 µs +15.6%
✅ gather/large_f16_threads=1-internal/131072 16.99 µs 19.41 µs +14.2%
✅ logit_processing/seven_processor_chain_per_step 372.57 µs 420.23 µs +12.8%
✅ qwen3_sampling_processors/top_k_top_p_fast 751.00 µs 845.42 µs +12.6%
✅ sampling_latency/top_k_per_token 63.49 µs 70.68 µs +11.3%
✅ qwen3_sampling_processors/top_k_top_p_full_sort_baseline 7.28 ms 8.09 ms +11.1%
✅ qwen3_sampling_processors/top_k_partial_selection 162.56 µs 180.49 µs +11.0%
✅ gather/medium_f32_threads=1-internal/32768 4.53 µs 4.91 µs +8.4%
✅ matmul/medium_generic_f16_threads=8/32x512x512 47.35 µs 51.24 µs +8.2%
✅ qwen3_sampling_processors/top_p_fast_after_top_k 586.57 µs 628.44 µs +7.1%
✅ reduce_mean/small_f32_threads=1-internal/4096 16.80 µs 17.78 µs +5.8%
✅ add/large_bf16_threads=1-internal/4194304 1.86 ms 1.96 ms +5.4%
✅ matmul/small_generic_f16_threads=8/1x256x256 40.08 µs 42.16 µs +5.2%
✅ add/small_f16_threads=1-internal/1024 519.3 ns 543.0 ns +4.6%
✅ gather/large_bf16_threads=1-internal/131072 18.47 µs 19.19 µs +3.9%
✅ add/medium_f16_threads=1-internal/262144 118.61 µs 121.55 µs +2.5%
✅ matmul/medium_generic_bf16_threads=8/32x512x512 660.49 µs 675.94 µs +2.3%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 7.09 ms 7.24 ms +2.2%
✅ matmul/large_generic_bf16_threads=1/32x1024x1024 2.40 ms 2.45 ms +2.1%
✅ gather/large_f32_threads=1-internal/131072 52.28 µs 53.25 µs +1.8%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.73 ms 2.78 ms +1.8%
✅ block_quantized_matmul_cached_dense/mxfp4_cached_dense_repeated_call/1x1024x1024 139.74 µs 142.19 µs +1.8%
✅ block_quantized_matmul_cached_dense/mxfp4_uncached_dequant_each_call/1x1024x1024 1.37 ms 1.39 ms +1.5%
✅ reduce_mean/medium_f32_threads=1-internal/65536 277.86 µs 281.89 µs +1.5%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 11.04 ms 11.19 ms +1.4%
✅ gather/medium_f16_threads=1-internal/32768 2.81 µs 2.84 µs +1.2%
✅ gather/small_f32_threads=1-internal/4096 757.6 ns 765.6 ns +1.1%
✅ block_quantized_moe_cached_dense/mxfp4_uncached_expert_dequant_each_call/rows=1,H=256,I=256,E=4,top_k=1 588.09 µs 593.02 µs +0.8%
✅ tokenization/decode_tokens_per_second 7.54 ms 7.60 ms +0.8%
✅ matmul/small_generic_f16_threads=1/1x256x256 39.11 µs 39.40 µs +0.7%
✅ matmul/large_generic_f16_threads=8/32x1024x1024 120.87 µs 121.73 µs +0.7%
✅ tokenization/encode_tokens_per_second 445.81 µs 447.93 µs +0.5%
✅ matmul/small_generic_bf16_threads=8/1x256x256 41.73 µs 41.85 µs +0.3%
✅ matmul/medium_generic_f32_threads=8/32x512x512 1.85 ms 1.85 ms +0.2%
✅ gather/small_bf16_threads=1-internal/4096 535.2 ns 534.3 ns -0.2%
✅ matmul/small_generic_f32_threads=1/1x256x256 48.20 µs 48.08 µs -0.3%
✅ matmul/medium_generic_f16_threads=1/32x512x512 40.40 µs 40.21 µs -0.5%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 97.05 µs 96.54 µs -0.5%
✅ add/large_f16_threads=1-internal/4194304 1.89 ms 1.88 ms -0.6%
✅ matmul/large_generic_bf16_threads=8/32x1024x1024 2.40 ms 2.37 ms -1.1%
✅ add/small_f32_threads=1-internal/1024 239.5 ns 236.5 ns -1.3%
✅ gather/small_f16_threads=1-internal/4096 557.6 ns 549.2 ns -1.5%
✅ block_quantized_matmul_cached_dense/mxfp4_preexpanded_dense_oncelock_like_proxy/1x1024x1024 120.54 µs 118.69 µs -1.5%
✅ add/large_f32_threads=1-internal/4194304 971.95 µs 955.44 µs -1.7%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 660.79 µs 647.35 µs -2.0%
✅ matmul/small_generic_bf16_threads=1/1x256x256 40.32 µs 39.45 µs -2.2%
✅ block_quantized_moe_cached_dense/mxfp4_cached_dense_expert_repeated_call/rows=1,H=256,I=256,E=4,top_k=1 121.74 µs 118.50 µs -2.7%
✅ gather/medium_bf16_threads=1-internal/32768 2.94 µs 2.86 µs -2.7%
✅ sampling_latency/greedy_per_token 3.72 µs 3.59 µs -3.6%
✅ matmul/small_generic_f32_threads=8/1x256x256 59.04 µs 56.27 µs -4.7%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.19 ms 1.12 ms -5.8%
✅ add/medium_bf16_threads=1-internal/262144 128.21 µs 116.91 µs -8.8%
🟢 add/small_bf16_threads=1-internal/1024 632.8 ns 512.5 ns -19.0%
🟢 add/medium_f32_threads=1-internal/262144 36.39 µs 28.63 µs -21.3%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 5.28 4.61 5.70 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants