Repository navigation
chore(squad): scribe round 9 — 27B decode roofline profile + Scan-capture scoping - #556
Merged
Merged
Conversation
…ture scoping consolidation Consolidate 3 decision notes into decisions.md (20332->20465 B, under 20480 gate): - cohaagen 27B decode profile: Scan (LinearAttention recurrence) = 56.5%, ~35x off HBM roofline, structural (control-flow un-capturable), not kernel-bound. - mary Scan-capture scoping: single-trip Scan->capture is NOT a contained increment (prefill/decode share one plan); needs runtime shape-conditional dual-path infra. Recorded as PENDING-JUSTIN decision (high blast radius in #443/#543 core). - mary #554 session-reuse recurrent-state reset (merged). Archived round-8 wave record to decisions-archive/2026-07.md; compacted cohaagen history (archive split); appended mary/cohaagen histories; wrote orchestration + session logs. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
justinchuby
enabled auto-merge (squash)
July 31, 2026 10:41
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #556 +/- ##
=======================================
Coverage 80.59% 80.59%
=======================================
Files 315 315
Lines 123446 123446
Branches 123446 123446
=======================================
Hits 99489 99489
Misses 19907 19907
Partials 4050 4050
Flags with carried forward coverage won't be shown. Click here to find out more. 🚀 New features to boost your workflow:
|
|
| Status | Scenario | Base | PR | Change |
|---|---|---|---|---|
matmul/small_generic_f32_threads=8/1x256x256 |
52.48 µs | 67.68 µs | +28.9% | |
matmul/medium_generic_f16_threads=8/32x512x512 |
33.38 µs | 43.03 µs | +28.9% | |
matmul/medium_generic_bf16_threads=8/32x512x512 |
492.36 µs | 589.44 µs | +19.7% | |
gather/large_bf16_threads=1-internal/131072 |
20.97 µs | 24.16 µs | +15.2% | |
| ✅ | matmul/medium_generic_bf16_threads=1/32x512x512 |
577.47 µs | 625.61 µs | +8.3% |
| ✅ | gather/medium_f32_threads=1-internal/32768 |
5.46 µs | 5.84 µs | +6.8% |
| ✅ | matmul/large_generic_f32_threads=1/32x1024x1024 |
10.72 ms | 11.37 ms | +6.0% |
| ✅ | matmul/small_generic_f16_threads=1/1x256x256 |
37.60 µs | 39.75 µs | +5.7% |
| ✅ | matmul/small_generic_f32_threads=1/1x256x256 |
53.80 µs | 56.22 µs | +4.5% |
| ✅ | matmul/small_generic_f16_threads=8/1x256x256 |
41.04 µs | 41.82 µs | +1.9% |
| ✅ | gather/medium_bf16_threads=1-internal/32768 |
3.27 µs | 3.32 µs | +1.6% |
| ✅ | add/medium_bf16_threads=1-internal/262144 |
3.12 ms | 3.03 ms | -2.8% |
| ✅ | tokenization/encode_tokens_per_second |
451.90 µs | 425.90 µs | -5.8% |
| ✅ | add/small_f32_threads=1-internal/1024 |
214.1 ns | 201.3 ns | -6.0% |
| ✅ | matmul/small_generic_bf16_threads=1/1x256x256 |
46.52 µs | 43.25 µs | -7.0% |
| ✅ | matmul/medium_generic_f16_threads=1/32x512x512 |
37.00 µs | 34.34 µs | -7.2% |
| ✅ | gather/medium_f16_threads=1-internal/32768 |
3.81 µs | 3.53 µs | -7.2% |
| ✅ | gather/small_f16_threads=1-internal/4096 |
597.5 ns | 553.1 ns | -7.4% |
| ✅ | add/small_bf16_threads=1-internal/1024 |
15.97 µs | 14.77 µs | -7.5% |
| ✅ | gather/small_bf16_threads=1-internal/4096 |
606.9 ns | 558.5 ns | -8.0% |
| ✅ | matmul/large_generic_f16_threads=1/32x1024x1024 |
108.06 µs | 95.93 µs | -11.2% |
| ✅ | gather/large_f16_threads=1-internal/131072 |
22.29 µs | 19.78 µs | -11.3% |
| ✅ | matmul/small_generic_bf16_threads=8/1x256x256 |
50.63 µs | 44.34 µs | -12.4% |
| ✅ | kv_cache/alloc_dealloc_pages |
42.35 µs | 36.60 µs | -13.6% |
| ✅ | add/medium_f32_threads=1-internal/262144 |
3.19 ms | 2.76 ms | -13.6% |
| ✅ | logit_processing/seven_processor_chain_per_step |
1.45 ms | 1.25 ms | -13.8% |
| ✅ | sampling_latency/min_p_per_token |
395.76 µs | 337.01 µs | -14.8% |
| 🟢 | grammar_masking/llguidance_compute_mask/32 |
90.07 µs | 75.02 µs | -16.7% |
| 🟢 | reduce_mean/large_f32_threads=1-internal/262144 |
1.22 ms | 1.01 ms | -16.8% |
| 🟢 | sampling_latency/top_p_per_token |
1.22 ms | 1.01 ms | -17.1% |
| 🟢 | add/medium_f16_threads=1-internal/262144 |
3.38 ms | 2.79 ms | -17.3% |
| 🟢 | tokenization/decode_tokens_per_second |
7.75 ms | 6.33 ms | -18.3% |
| 🟢 | add/large_f32_threads=1-internal/4194304 |
55.59 ms | 44.79 ms | -19.4% |
| 🟢 | sampling_latency/greedy_per_token |
4.00 µs | 3.21 µs | -19.9% |
| 🟢 | matmul/large_generic_f16_threads=8/32x1024x1024 |
136.94 µs | 108.03 µs | -21.1% |
| 🟢 | gather/small_f32_threads=1-internal/4096 |
892.7 ns | 689.3 ns | -22.8% |
| 🟢 | add/small_f16_threads=1-internal/1024 |
17.42 µs | 13.29 µs | -23.7% |
| 🟢 | sampling_latency/top_k_per_token |
621.62 µs | 474.02 µs | -23.7% |
| 🟢 | matmul/large_generic_bf16_threads=1/32x1024x1024 |
2.62 ms | 1.99 ms | -24.1% |
| 🟢 | matmul/large_generic_f32_threads=8/32x1024x1024 |
7.05 ms | 5.27 ms | -25.3% |
| 🟢 | matmul/large_generic_bf16_threads=8/32x1024x1024 |
1.89 ms | 1.41 ms | -25.3% |
| 🟢 | reduce_mean/medium_f32_threads=1-internal/65536 |
337.39 µs | 249.29 µs | -26.1% |
| 🟢 | add/large_f16_threads=1-internal/4194304 |
59.86 ms | 43.58 ms | -27.2% |
| 🟢 | gather/large_f32_threads=1-internal/131072 |
54.29 µs | 39.23 µs | -27.7% |
| 🟢 | matmul/medium_generic_f32_threads=8/32x512x512 |
1.69 ms | 1.19 ms | -29.5% |
| 🟢 | add/large_bf16_threads=1-internal/4194304 |
64.29 ms | 44.66 ms | -30.5% |
| 🟢 | reduce_mean/small_f32_threads=1-internal/4096 |
23.79 µs | 15.52 µs | -34.8% |
| 🟢 | matmul/medium_generic_f32_threads=1/32x512x512 |
4.40 ms | 2.59 ms | -41.1% |
Visual flags:
Host info
CPU: Apple M1 (Virtual)
Cores: 3
OS: Darwin 25.5.0 arm64
Rust: rustc 1.97.1 (8bab26f4f 2026-07-14)
Load avg: { 2.97 4.16 7.75 }
What this cannot catch
- Regressions in code paths not covered by these benchmarks (e.g., end-to-end decode with a real model)
- Sub-threshold regressions that compound over multiple PRs
- Performance changes that only manifest under GPU execution
- Latency changes in the ORT integration path (these benchmarks exercise the native Rust kernels)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
State consolidation only (no code). Merges 3 decision notes into decisions.md (20332→20465 B, under gate).
Key durable record: the ⚑ PENDING-JUSTIN 27B Scan→CUDA-capture decision — the ~15-30× decode lever is NOT a contained increment (prefill/decode share one plan/session; static single-trip inline corrupts prefill). Correct fix = runtime shape-conditional dual-path capture infra, which reshapes the delicate #443/#543 control-flow capture core. Awaiting @justinchuby go-ahead before touching that core.
Also records: 27B decode profile (Scan=56.5%, ~35× off HBM roofline, structural not kernel), and the #554 session-reuse fix. Round-8 wave record archived; histories compacted.