Skip to content

chore(decisions): consolidate GAP-3 native-paged-decode decisions (#565-#568) - #569

Merged
justinchuby merged 1 commit into
mainfrom
chore/scribe-gap3-consolidation
Jul 31, 2026
Merged

justinchuby merged 1 commit into
mainfrom
chore/scribe-gap3-consolidation

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

Consolidates the accumulated GAP-3 native-paged-decode decision inbox (7 notes) into the canonical .squad/decisions.md, clearing the drop-box.

Folds in the decisions from the 4 merged GAP-3 PRs plus the Scan-capture deferral:

decisions.md 20465 → 28227 bytes (under the 30KB healthy threshold; no archiving needed). Inbox now README-only. Docs-only change under .squad/.

Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com

#568)

Consolidated 7 inbox decisions from GAP-3 increments (Inc-A/C/D/D.1 pipeline construction, paging, device-KV read-out, f16 support) and related work (Inc-C test-rigor fix, Scan-capture deferral decision, Scan 1a single-trip inline). Merged entries capture present-KV threading mechanics, f16/f32 geometry proof, and deferred device-graph-registry refactor.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby enabled auto-merge (squash) July 31, 2026 23:49
@codecov

codecov Bot commented Jul 31, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 81.32%. Comparing base (c1a2450) to head (86d9c75).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main     #569      +/-   ##
==========================================
+ Coverage   80.60%   81.32%   +0.72%     
==========================================
  Files         315      315              
  Lines      123574   123574              
  Branches   123574   123574              
==========================================
+ Hits        99606   100501     +895     
+ Misses      19916    19015     -901     
- Partials     4052     4058       +6     
Flag Coverage Δ
cli-ort-linux 83.27% <ø> (ø)
cli-ort-windows 82.67% <ø> (-0.11%) ⬇️
mlas 78.72% <ø> (+0.81%) ⬆️
offline 81.28% <ø> (+0.75%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.
see 7 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@justinchuby
justinchuby merged commit f3c6e79 into main Jul 31, 2026
14 checks passed
@justinchuby
justinchuby deleted the chore/scribe-gap3-consolidation branch July 31, 2026 23:57
@github-actions

github-actions Bot commented Aug 1, 2026

Copy link
Copy Markdown

🔴 Benchmark Regression Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
🔴 matmul/medium_generic_bf16_threads=8/32x512x512 631.13 µs 1.36 ms +116.0%
🔴 matmul/large_generic_f32_threads=1/32x1024x1024 10.64 ms 16.69 ms +56.9%
🔴 matmul/medium_generic_f16_threads=8/32x512x512 49.80 µs 73.53 µs +47.7%
🔴 logit_processing/seven_processor_chain_per_step 1.34 ms 1.84 ms +37.7%
🔴 matmul/large_generic_f16_threads=8/32x1024x1024 110.71 µs 150.40 µs +35.8%
🔴 matmul/medium_generic_bf16_threads=1/32x512x512 591.66 µs 801.43 µs +35.5%
🔴 sampling_latency/top_p_per_token 1.09 ms 1.44 ms +31.9%
🔴 matmul/large_generic_bf16_threads=1/32x1024x1024 2.36 ms 3.09 ms +31.3%
⚠️ matmul/large_generic_f16_threads=1/32x1024x1024 98.78 µs 128.32 µs +29.9%
⚠️ grammar_masking/llguidance_compute_mask/32 84.84 µs 107.70 µs +26.9%
⚠️ tokenization/encode_tokens_per_second 440.29 µs 557.57 µs +26.6%
⚠️ tokenization/decode_tokens_per_second 6.94 ms 8.71 ms +25.6%
⚠️ sampling_latency/top_k_per_token 550.94 µs 681.97 µs +23.8%
⚠️ kv_cache/alloc_dealloc_pages 40.51 µs 49.62 µs +22.5%
⚠️ sampling_latency/greedy_per_token 3.55 µs 4.27 µs +20.3%
⚠️ sampling_latency/min_p_per_token 379.56 µs 447.11 µs +17.8%
✅ matmul/large_generic_f32_threads=8/32x1024x1024 6.18 ms 6.93 ms +12.1%
✅ matmul/small_generic_bf16_threads=8/1x256x256 39.17 µs 43.23 µs +10.4%
✅ add/large_f32_threads=1-internal/4194304 45.80 ms 48.43 ms +5.7%
✅ matmul/medium_generic_f16_threads=1/32x512x512 40.07 µs 41.25 µs +2.9%
✅ matmul/large_generic_bf16_threads=8/32x1024x1024 2.19 ms 2.23 ms +2.1%
✅ add/medium_f16_threads=1-internal/262144 2.88 ms 2.93 ms +1.9%
✅ add/small_f32_threads=1-internal/1024 215.3 ns 216.8 ns +0.7%
✅ gather/small_f16_threads=1-internal/4096 545.2 ns 549.0 ns +0.7%
✅ matmul/medium_generic_f32_threads=1/32x512x512 2.73 ms 2.74 ms +0.3%
✅ gather/large_f16_threads=1-internal/131072 19.75 µs 19.77 µs +0.1%
✅ matmul/small_generic_f16_threads=1/1x256x256 40.47 µs 40.50 µs +0.1%
✅ add/small_f16_threads=1-internal/1024 14.25 µs 14.19 µs -0.4%
✅ matmul/small_generic_f32_threads=1/1x256x256 51.41 µs 51.16 µs -0.5%
✅ add/large_bf16_threads=1-internal/4194304 48.40 ms 47.70 ms -1.5%
✅ add/medium_bf16_threads=1-internal/262144 3.05 ms 2.98 ms -2.1%
✅ add/small_bf16_threads=1-internal/1024 14.83 µs 14.35 µs -3.3%
✅ add/medium_f32_threads=1-internal/262144 2.95 ms 2.85 ms -3.3%
✅ add/large_f16_threads=1-internal/4194304 49.57 ms 47.31 ms -4.6%
✅ reduce_mean/large_f32_threads=1-internal/262144 1.18 ms 1.10 ms -6.2%
✅ matmul/medium_generic_f32_threads=8/32x512x512 1.76 ms 1.65 ms -6.3%
✅ gather/small_bf16_threads=1-internal/4096 582.9 ns 545.9 ns -6.4%
✅ reduce_mean/medium_f32_threads=1-internal/65536 292.42 µs 269.52 µs -7.8%
✅ gather/medium_bf16_threads=1-internal/32768 3.15 µs 2.90 µs -7.9%
✅ matmul/small_generic_bf16_threads=1/1x256x256 43.94 µs 40.12 µs -8.7%
✅ gather/large_bf16_threads=1-internal/131072 19.17 µs 17.46 µs -8.9%
✅ matmul/small_generic_f16_threads=8/1x256x256 44.54 µs 40.33 µs -9.5%
✅ gather/small_f32_threads=1-internal/4096 807.1 ns 721.7 ns -10.6%
✅ matmul/small_generic_f32_threads=8/1x256x256 62.00 µs 55.41 µs -10.6%
✅ reduce_mean/small_f32_threads=1-internal/4096 18.96 µs 16.75 µs -11.6%
🟢 gather/medium_f32_threads=1-internal/32768 5.59 µs 4.60 µs -17.7%
🟢 gather/medium_f16_threads=1-internal/32768 3.65 µs 2.99 µs -18.0%
🟢 gather/large_f32_threads=1-internal/131072 63.44 µs 39.35 µs -38.0%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.4.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 5.72 5.68 7.10 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant