Skip to content

chore(squad): scribe round 9 — 27B decode roofline profile + Scan-capture scoping - #556

Merged
justinchuby merged 1 commit into
mainfrom
chore/scribe-round9-27b-profile
Jul 31, 2026
Merged

justinchuby merged 1 commit into
mainfrom
chore/scribe-round9-27b-profile

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

State consolidation only (no code). Merges 3 decision notes into decisions.md (20332→20465 B, under gate).

Key durable record: the ⚑ PENDING-JUSTIN 27B Scan→CUDA-capture decision — the ~15-30× decode lever is NOT a contained increment (prefill/decode share one plan/session; static single-trip inline corrupts prefill). Correct fix = runtime shape-conditional dual-path capture infra, which reshapes the delicate #443/#543 control-flow capture core. Awaiting @justinchuby go-ahead before touching that core.

Also records: 27B decode profile (Scan=56.5%, ~35× off HBM roofline, structural not kernel), and the #554 session-reuse fix. Round-8 wave record archived; histories compacted.

…ture scoping consolidation

Consolidate 3 decision notes into decisions.md (20332->20465 B, under 20480 gate):
- cohaagen 27B decode profile: Scan (LinearAttention recurrence) = 56.5%, ~35x off
  HBM roofline, structural (control-flow un-capturable), not kernel-bound.
- mary Scan-capture scoping: single-trip Scan->capture is NOT a contained increment
  (prefill/decode share one plan); needs runtime shape-conditional dual-path infra.
  Recorded as PENDING-JUSTIN decision (high blast radius in #443/#543 core).
- mary #554 session-reuse recurrent-state reset (merged).

Archived round-8 wave record to decisions-archive/2026-07.md; compacted cohaagen
history (archive split); appended mary/cohaagen histories; wrote orchestration + session logs.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@justinchuby
justinchuby enabled auto-merge (squash) July 31, 2026 10:41
@codecov

codecov Bot commented Jul 31, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.59%. Comparing base (bcdae6f) to head (3d0402e).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@           Coverage Diff           @@
##             main     #556   +/-   ##
=======================================
  Coverage   80.59%   80.59%           
=======================================
  Files         315      315           
  Lines      123446   123446           
  Branches   123446   123446           
=======================================
  Hits        99489    99489           
  Misses      19907    19907           
  Partials     4050     4050           
Flag Coverage Δ
cli-ort-linux 83.27% <ø> (ø)
cli-ort-windows 82.78% <ø> (ø)
mlas 77.91% <ø> (ø)
offline 80.51% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@justinchuby
justinchuby merged commit 209b830 into main Jul 31, 2026
14 checks passed
@justinchuby
justinchuby deleted the chore/scribe-round9-27b-profile branch July 31, 2026 10:50
@github-actions

Copy link
Copy Markdown

⚠️ Benchmark Change Detected

Comparison of criterion micro-benchmarks: PR head vs merge-base, measured on the same runner in the same job (base first → PR second).

ℹ️ Absolute times are informational only — they vary with runner load. The % change column is the reliable signal because both sides ran under identical conditions.

Status Scenario Base PR Change
⚠️ matmul/small_generic_f32_threads=8/1x256x256 52.48 µs 67.68 µs +28.9%
⚠️ matmul/medium_generic_f16_threads=8/32x512x512 33.38 µs 43.03 µs +28.9%
⚠️ matmul/medium_generic_bf16_threads=8/32x512x512 492.36 µs 589.44 µs +19.7%
⚠️ gather/large_bf16_threads=1-internal/131072 20.97 µs 24.16 µs +15.2%
✅ matmul/medium_generic_bf16_threads=1/32x512x512 577.47 µs 625.61 µs +8.3%
✅ gather/medium_f32_threads=1-internal/32768 5.46 µs 5.84 µs +6.8%
✅ matmul/large_generic_f32_threads=1/32x1024x1024 10.72 ms 11.37 ms +6.0%
✅ matmul/small_generic_f16_threads=1/1x256x256 37.60 µs 39.75 µs +5.7%
✅ matmul/small_generic_f32_threads=1/1x256x256 53.80 µs 56.22 µs +4.5%
✅ matmul/small_generic_f16_threads=8/1x256x256 41.04 µs 41.82 µs +1.9%
✅ gather/medium_bf16_threads=1-internal/32768 3.27 µs 3.32 µs +1.6%
✅ add/medium_bf16_threads=1-internal/262144 3.12 ms 3.03 ms -2.8%
✅ tokenization/encode_tokens_per_second 451.90 µs 425.90 µs -5.8%
✅ add/small_f32_threads=1-internal/1024 214.1 ns 201.3 ns -6.0%
✅ matmul/small_generic_bf16_threads=1/1x256x256 46.52 µs 43.25 µs -7.0%
✅ matmul/medium_generic_f16_threads=1/32x512x512 37.00 µs 34.34 µs -7.2%
✅ gather/medium_f16_threads=1-internal/32768 3.81 µs 3.53 µs -7.2%
✅ gather/small_f16_threads=1-internal/4096 597.5 ns 553.1 ns -7.4%
✅ add/small_bf16_threads=1-internal/1024 15.97 µs 14.77 µs -7.5%
✅ gather/small_bf16_threads=1-internal/4096 606.9 ns 558.5 ns -8.0%
✅ matmul/large_generic_f16_threads=1/32x1024x1024 108.06 µs 95.93 µs -11.2%
✅ gather/large_f16_threads=1-internal/131072 22.29 µs 19.78 µs -11.3%
✅ matmul/small_generic_bf16_threads=8/1x256x256 50.63 µs 44.34 µs -12.4%
✅ kv_cache/alloc_dealloc_pages 42.35 µs 36.60 µs -13.6%
✅ add/medium_f32_threads=1-internal/262144 3.19 ms 2.76 ms -13.6%
✅ logit_processing/seven_processor_chain_per_step 1.45 ms 1.25 ms -13.8%
✅ sampling_latency/min_p_per_token 395.76 µs 337.01 µs -14.8%
🟢 grammar_masking/llguidance_compute_mask/32 90.07 µs 75.02 µs -16.7%
🟢 reduce_mean/large_f32_threads=1-internal/262144 1.22 ms 1.01 ms -16.8%
🟢 sampling_latency/top_p_per_token 1.22 ms 1.01 ms -17.1%
🟢 add/medium_f16_threads=1-internal/262144 3.38 ms 2.79 ms -17.3%
🟢 tokenization/decode_tokens_per_second 7.75 ms 6.33 ms -18.3%
🟢 add/large_f32_threads=1-internal/4194304 55.59 ms 44.79 ms -19.4%
🟢 sampling_latency/greedy_per_token 4.00 µs 3.21 µs -19.9%
🟢 matmul/large_generic_f16_threads=8/32x1024x1024 136.94 µs 108.03 µs -21.1%
🟢 gather/small_f32_threads=1-internal/4096 892.7 ns 689.3 ns -22.8%
🟢 add/small_f16_threads=1-internal/1024 17.42 µs 13.29 µs -23.7%
🟢 sampling_latency/top_k_per_token 621.62 µs 474.02 µs -23.7%
🟢 matmul/large_generic_bf16_threads=1/32x1024x1024 2.62 ms 1.99 ms -24.1%
🟢 matmul/large_generic_f32_threads=8/32x1024x1024 7.05 ms 5.27 ms -25.3%
🟢 matmul/large_generic_bf16_threads=8/32x1024x1024 1.89 ms 1.41 ms -25.3%
🟢 reduce_mean/medium_f32_threads=1-internal/65536 337.39 µs 249.29 µs -26.1%
🟢 add/large_f16_threads=1-internal/4194304 59.86 ms 43.58 ms -27.2%
🟢 gather/large_f32_threads=1-internal/131072 54.29 µs 39.23 µs -27.7%
🟢 matmul/medium_generic_f32_threads=8/32x512x512 1.69 ms 1.19 ms -29.5%
🟢 add/large_bf16_threads=1-internal/4194304 64.29 ms 44.66 ms -30.5%
🟢 reduce_mean/small_f32_threads=1-internal/4096 23.79 µs 15.52 µs -34.8%
🟢 matmul/medium_generic_f32_threads=1/32x512x512 4.40 ms 2.59 ms -41.1%

Visual flags: ⚠️ ≥ 15% slower, 🔴 ≥ 30% slower — calibrated against measured runner noise (~27% worst-case on multi-threaded matmul)

Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 2.97 4.16 7.75 }
What this cannot catch
  • Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
  • Sub-threshold regressions that compound over multiple PRs
  • Performance changes that only manifest under GPU execution
  • Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants