Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 15 additions & 0 deletions docs/audit/backend/rocm/ROCM_AUDIT.md
Original file line number Diff line number Diff line change
Expand Up @@ -341,6 +341,21 @@ become the on-silicon **oracle** the compiled path validates against.
GPU-free codegen gate. The CPU half (AVX-512) landed in the x86 backend
(`avx512_reduce_f32.cpp`) — so the reduction family now has a real optimized
kernel on both devices we have hardware for. Status `compiled`.
- **Elementwise unary math** (S2 scalar-math / stability family, 2026-06-25):
a `tessera_rocm.unary` directive + `generate-rocm-unary-kernel` pass emitting
a flat per-element kernel (one thread per element), the unary sibling of the
activation lane. Covers `exp`/`log`/`sqrt`/`rsqrt`/`reciprocal`/`abs`
(`absolute`)/`sign`/`erf`/`tanh`/`sigmoid`/`log1p`/`expm1`/`softplus`
(softplus stable: `log1p(exp(-|x|)) + max(x,0)`); transcendentals lower
through the `math` → ROCDL path. New `runtime.launch()` lane
`rocm_unary_compiled`, dispatched by op name; f16/bf16/f32 storage, f32
compute. Validated on gfx1151 vs numpy across kind × dtype × shape incl.
rank-3 (`test_rocm_unary_compiled.py`) + a GPU-free codegen gate. The CPU
half landed in the x86 backend as an AVX-512 kernel for the **algebraic
subset** (`sqrt`/`rsqrt`/`reciprocal`/`abs`/`neg`/`sign`, direct intrinsics,
no polynomial approx — `avx512_unary_f32.cpp`, validated standalone); the
transcendentals stay numpy-reference on CPU (no fused x86 claim). Status
`compiled`.
- **rmsnorm / layer_norm** (2026-06-25): the row-reduction siblings of the
softmax kernel — a `tessera_rocm.norm` directive + `generate-rocm-norm-kernel`
pass (one workgroup per row). rmsnorm tree-reduces Σx² in one pass; layer_norm
Expand Down
2 changes: 2 additions & 0 deletions docs/audit/generated/runtime_abi.csv
Original file line number Diff line number Diff line change
Expand Up @@ -578,8 +578,10 @@ x86,tessera_x86_amx_gemm_s8s8_s32,amx_gemm_s8s8_s32,,src/compiler/codegen/tesser
x86,tessera_x86_amx_gemm_s8s8_s32,amx_gemm_s8s8_s32,,src/compiler/codegen/tessera_x86_backend/src/kernels/amx_gemm_int8.cpp
x86,tessera_x86_avx512_gemm_bf16,avx512_gemm,bf16,src/compiler/codegen/tessera_x86_backend/src/kernels/avx512_gemm_bf16.cpp
x86,tessera_x86_avx512_reduce_f32,avx512_reduce,f32,src/compiler/codegen/tessera_x86_backend/src/kernels/avx512_reduce_f32.cpp
x86,tessera_x86_avx512_unary_f32,avx512_unary,f32,src/compiler/codegen/tessera_x86_backend/src/kernels/avx512_unary_f32.cpp
x86,tessera_x86_avx512_vnni_gemm_u8s8_s32,avx512_vnni_gemm_u8s8_s32,,src/compiler/codegen/tessera_x86_backend/src/kernels/avx512_vnni_gemm_int8.cpp
x86,tessera_x86_epilogue_bias_fp32,epilogue_bias_fp32,,src/compiler/codegen/tessera_x86_backend/src/kernels/epilogue.cpp
x86,tessera_x86_epilogue_bias_gelu_fp32,epilogue_bias_gelu_fp32,,src/compiler/codegen/tessera_x86_backend/src/kernels/epilogue.cpp
x86,tessera_x86_reference_gemm_bf16,reference_gemm,bf16,src/compiler/codegen/tessera_x86_backend/src/kernels/avx512_gemm_bf16.cpp
x86,tessera_x86_reference_reduce_f32,reference_reduce,f32,src/compiler/codegen/tessera_x86_backend/src/kernels/avx512_reduce_f32.cpp
x86,tessera_x86_reference_unary_f32,reference_unary,f32,src/compiler/codegen/tessera_x86_backend/src/kernels/avx512_unary_f32.cpp
4 changes: 2 additions & 2 deletions docs/audit/generated/runtime_abi.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ Human-readable view. The canonical machine-readable artifact is `runtime_abi.csv

## Headline

- **328** unique `extern "C" tessera_*` C ABI symbols across all backends.
- **330** unique `extern "C" tessera_*` C ABI symbols across all backends.
- **6 / 6** core runtime headers present.
- **134** Apple GPU kernel families with per-dtype variants.

Expand All @@ -26,7 +26,7 @@ Human-readable view. The canonical machine-readable artifact is `runtime_abi.csv
| `apple` | 304 |
| `nvidia` | 4 |
| `rocm` | 10 |
| `x86` | 10 |
| `x86` | 12 |

## Apple GPU kernel families × dtype matrix

Expand Down
1 change: 1 addition & 0 deletions docs/audit/generated/runtime_execution_matrix.csv
Original file line number Diff line number Diff line change
Expand Up @@ -19,4 +19,5 @@ rocm,rocm_reduce_compiled,native_gpu,1,rocm_reduce_compiled,success,hip_runtime,
rocm,rocm_rope_compiled,native_gpu,1,rocm_rope_compiled,success,hip_runtime,"ROCm rope artifact runs the COMPILER-GENERATED interleaved-pair rotary-position-embedding kernel (one workgroup per row): tessera-opt generates + serializes the kernel to hsaco in-process, then HIP loads + launches it."
rocm,rocm_silu_mul_compiled,native_gpu,1,rocm_silu_mul_compiled,success,hip_runtime,"ROCm silu_mul artifact runs the COMPILER-GENERATED flat 2-operand elementwise SwiGLU gate-multiply silu(a)·b (one thread per element): tessera-opt generates + serializes the kernel to hsaco in-process, then HIP loads + launches it."
rocm,rocm_softmax_compiled,native_gpu,1,rocm_softmax_compiled,success,hip_runtime,"ROCm softmax artifact runs the COMPILER-GENERATED RDNA row-reduction kernel (stable softmax over the last axis, one workgroup per row, LDS tree-reduce): tessera-opt generates + serializes the kernel to hsaco in-process, then HIP loads + launches it. The first non-matmul/non-WMMA compiled ROCm kernel."
rocm,rocm_unary_compiled,native_gpu,1,rocm_unary_compiled,success,hip_runtime,"ROCm unary artifact runs the COMPILER-GENERATED flat elementwise unary-math kernel (S2 scalar-math/stability: exp/log/sqrt/rsqrt/reciprocal/abs/sign/erf/tanh/sigmoid/log1p/expm1/softplus, one thread per element): tessera-opt generates + serializes the kernel to hsaco in-process, then HIP loads + launches it. Dispatched by op name."
rocm,rocm_wmma,native_gpu,1,rocm_wmma,success,hip_runtime,"ROCm matmul via the hand-written RDNA WMMA GEMM (tessera_rocm_wmma_gemm_{f16,bf16} C ABI symbol, HIPRTC-compiled for the device arch). Now the reference ORACLE + availability fallback for the compiled lane (rocm_compiled) — still directly selectable by stamping compiler_path=""rocm_wmma""."
2 changes: 2 additions & 0 deletions docs/audit/generated/runtime_execution_matrix.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,7 @@ Single source of truth for what `runtime.launch()` does with each `(target, comp
| `rocm` | `rocm_rope_compiled` | `rocm_rope_compiled` | `native_gpu` | `hip_runtime` | ROCm rope artifact runs the COMPILER-GENERATED interleaved-pair rotary-position-embedding kernel (one workgroup per row): tessera-opt generates + serializes the kernel to hsaco in-process, then HIP loads + launches it. |
| `rocm` | `rocm_silu_mul_compiled` | `rocm_silu_mul_compiled` | `native_gpu` | `hip_runtime` | ROCm silu_mul artifact runs the COMPILER-GENERATED flat 2-operand elementwise SwiGLU gate-multiply silu(a)·b (one thread per element): tessera-opt generates + serializes the kernel to hsaco in-process, then HIP loads + launches it. |
| `rocm` | `rocm_softmax_compiled` | `rocm_softmax_compiled` | `native_gpu` | `hip_runtime` | ROCm softmax artifact runs the COMPILER-GENERATED RDNA row-reduction kernel (stable softmax over the last axis, one workgroup per row, LDS tree-reduce): tessera-opt generates + serializes the kernel to hsaco in-process, then HIP loads + launches it. The first non-matmul/non-WMMA compiled ROCm kernel. |
| `rocm` | `rocm_unary_compiled` | `rocm_unary_compiled` | `native_gpu` | `hip_runtime` | ROCm unary artifact runs the COMPILER-GENERATED flat elementwise unary-math kernel (S2 scalar-math/stability: exp/log/sqrt/rsqrt/reciprocal/abs/sign/erf/tanh/sigmoid/log1p/expm1/softplus, one thread per element): tessera-opt generates + serializes the kernel to hsaco in-process, then HIP loads + launches it. Dispatched by op name. |
| `rocm` | `rocm_wmma` | `rocm_wmma` | `native_gpu` | `hip_runtime` | ROCm matmul via the hand-written RDNA WMMA GEMM (tessera_rocm_wmma_gemm_{f16,bf16} C ABI symbol, HIPRTC-compiled for the device arch). Now the reference ORACLE + availability fallback for the compiled lane (rocm_compiled) — still directly selectable by stamping compiler_path="rocm_wmma". |

## Targets with no executable row
Expand Down Expand Up @@ -67,4 +68,5 @@ nvidia_sm80, nvidia_sm90, nvidia_sm100, rocm_gfx90a, rocm_gfx940, rocm_gfx942, r
| `rocm_rope_compiled` | AMD GPU RDNA rotary-position-embedding the Tessera compiler GENERATES (generate-rocm-rope-kernel -> ROCDL -> hsaco, in-process via tessera-opt), then HIP loads + launches it. Interleaved-pair RoPE over [M, D] (one workgroup per row); f32/f16/bf16 |
| `rocm_silu_mul_compiled` | AMD GPU RDNA SwiGLU gate-multiply the Tessera compiler GENERATES (generate-rocm-silu-mul-kernel -> ROCDL -> hsaco, in-process via tessera-opt), then HIP loads + launches it. Flat 2-operand elementwise silu(a)·b (one thread per element); the standalone analog of the fused SwiGLU gate-multiply; f32/f16/bf16 storage, f32 compute |
| `rocm_softmax_compiled` | AMD GPU RDNA row-reduction softmax the Tessera compiler GENERATES (generate-rocm-softmax-kernel -> ROCDL -> hsaco, in-process via tessera-opt), then HIP loads + launches it. Stable softmax over the last axis (one workgroup per row, LDS tree-reduce); the first non-matmul/non-WMMA compiled ROCm kernel. f32/f16/bf16 storage, f32 reduce |
| `rocm_unary_compiled` | AMD GPU RDNA flat elementwise unary-math kernel the Tessera compiler GENERATES (generate-rocm-unary-kernel -> ROCDL -> hsaco, in-process via tessera-opt), then HIP loads + launches it — the S2 scalar-math / stability family (exp/log/sqrt/rsqrt/reciprocal/abs/sign/erf/tanh/sigmoid/log1p/expm1/softplus), one thread per element, dispatched by op name; f32/f16/bf16 storage, f32 compute |
| `rocm_wmma` | AMD GPU RDNA WMMA matrix-core GEMM via the shipped libtessera_rocm_gemm.so tessera_rocm_wmma_gemm_{f16,bf16} C ABI symbol (HIPRTC-compiled for the device arch; f16/bf16 storage, f32 accumulate) |
30 changes: 15 additions & 15 deletions docs/audit/generated/test_coverage.csv
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
op,python_refs,lit_refs,negative_refs,total_refs,is_thinly_tested,dtype_variants,bucket,reason
abs,2,0,0,2,0,,directly_tested,2 direct test references
absolute,1,0,0,1,1,,structural_only,unclassified — defaults to structural_only
abs,3,0,0,3,0,bf16 f32,directly_tested,3 direct test references
absolute,2,0,0,2,0,bf16 f32,directly_tested,2 direct test references
acos,1,0,0,1,1,,covered_by_family,category default for 'elementwise'
adafactor,4,0,0,4,0,fp32 fp64,directly_tested,4 direct test references
adam,7,5,0,12,0,fp32 fp64,directly_tested,12 direct test references
Expand Down Expand Up @@ -171,11 +171,11 @@ ema_update,1,0,0,1,1,,structural_only,category default for 'grad_transform'
empty_state_tree,0,0,0,0,1,,structural_only,category default for 'state_tree'
eq,0,0,0,0,1,,covered_by_family,category default for 'comparison'
equiprob_band_partition,0,0,0,0,1,,structural_only,unclassified — defaults to structural_only
erf,1,0,0,1,1,,covered_by_family,category default for 'elementwise'
erf,2,0,0,2,0,bf16 f32,directly_tested,2 direct test references
erfc,1,0,0,1,1,,covered_by_family,category default for 'elementwise'
exp,4,0,0,4,0,,directly_tested,4 direct test references
exp,5,0,0,5,0,bf16 f32,directly_tested,5 direct test references
expand,1,0,0,1,1,,structural_only,unclassified — defaults to structural_only
expm1,1,0,0,1,1,,covered_by_family,category default for 'elementwise'
expm1,2,0,0,2,0,bf16 f32,directly_tested,2 direct test references
factorized_matmul,4,0,0,4,0,,directly_tested,4 direct test references
factorized_pos_emb,6,0,1,6,0,,directly_tested,6 direct test references
fake_quantize,3,0,0,3,0,,directly_tested,3 direct test references
Expand Down Expand Up @@ -247,8 +247,8 @@ lion,4,0,0,4,0,fp32 fp64,directly_tested,4 direct test references
load_balance_loss,2,0,0,2,0,,directly_tested,2 direct test references
load_sharded,3,0,1,3,0,,directly_tested,3 direct test references
load_state,14,0,2,14,0,,directly_tested,14 direct test references
log,3,0,0,3,0,,directly_tested,3 direct test references
log1p,2,0,0,2,0,,directly_tested,2 direct test references
log,4,0,0,4,0,bf16 f32,directly_tested,4 direct test references
log1p,3,0,0,3,0,bf16 f32,directly_tested,3 direct test references
log_cosh_loss,1,0,0,1,1,,covered_by_family,category default for 'loss'
log_softmax,5,5,0,10,0,,directly_tested,10 direct test references
logical_and,1,0,0,1,1,,structural_only,unclassified — defaults to structural_only
Expand Down Expand Up @@ -345,7 +345,7 @@ quantize_int8,3,0,0,3,0,,directly_tested,3 direct test references
quantize_nvfp4,11,0,2,11,0,fp16 fp32 fp4_e2m1 fp6_e2m3 fp8_e4m3 fp8_e5m2 int8 nvfp4,directly_tested,11 direct test references
quantized_matmul,5,0,0,5,0,f16,directly_tested,5 direct test references
rearrange,2,0,0,2,0,,directly_tested,2 direct test references
reciprocal,1,0,0,1,1,,structural_only,unclassified — defaults to structural_only
reciprocal,2,0,0,2,0,bf16 f32,directly_tested,2 direct test references
reduce,86,0,7,86,0,f32 fp16 fp32 fp4_e2m1 fp6_e2m3 fp8_e4m3 nvfp4,directly_tested,86 direct test references
reduce_scatter,0,4,0,4,0,,directly_tested,4 direct test references
relu,105,20,9,125,0,bf16 f16 f32 f64 fp32 fp8_e4m3 int8,directly_tested,125 direct test references
Expand Down Expand Up @@ -381,7 +381,7 @@ rope,19,10,0,29,0,bf16,directly_tested,29 direct test references
rope_merge,2,0,1,2,0,fp16 fp32,directly_tested,2 direct test references
rope_split,8,0,1,8,0,fp16 fp32,directly_tested,8 direct test references
round,1,0,0,1,1,,structural_only,unclassified — defaults to structural_only
rsqrt,2,0,0,2,0,,directly_tested,2 direct test references
rsqrt,3,0,0,3,0,bf16 f32,directly_tested,3 direct test references
safetensors_export,1,0,0,1,1,,structural_only,category default for 'aot'
save_sharded,1,0,1,1,1,,structural_only,category default for 'serialization'
save_state,6,0,2,6,0,,directly_tested,6 direct test references
Expand All @@ -399,9 +399,9 @@ seq2seq_loss,3,0,0,3,0,,directly_tested,3 direct test references
sgd,6,0,0,6,0,,directly_tested,6 direct test references
shard_map,0,0,0,0,1,,structural_only,category default for 'sharding'
sharded_dataset,2,0,0,2,0,,directly_tested,2 direct test references
sigmoid,4,0,0,4,0,bf16 f16 f32 fp32,directly_tested,4 direct test references
sigmoid,5,0,0,5,0,bf16 f16 f32 fp32,directly_tested,5 direct test references
sigmoid_safe,4,0,0,4,0,,directly_tested,4 direct test references
sign,2,0,0,2,0,,directly_tested,2 direct test references
sign,4,0,1,4,0,bf16 f32,directly_tested,4 direct test references
silu,84,2,5,86,0,bf16 f16 f32 f64 fp16 fp32 fp4_e2m1 fp6_e2m3 fp8_e4m3 nvfp4,directly_tested,86 direct test references
silu_mul,14,13,0,27,0,bf16 fp32,directly_tested,27 direct test references
simple_rnn_cell,5,0,0,5,0,,directly_tested,5 direct test references
Expand All @@ -410,17 +410,17 @@ sinh,1,0,0,1,1,,covered_by_family,category default for 'elementwise'
slice,11,0,0,11,0,,directly_tested,11 direct test references
smooth_l1_loss,1,0,0,1,1,,covered_by_family,category default for 'loss'
softcap,7,0,1,7,0,,directly_tested,7 direct test references
softmax,97,37,9,134,0,bf16 f16 f32 fp16 fp32 fp4_e2m1 fp6_e2m3 fp8_e4m3 nvfp4,directly_tested,134 direct test references
softmax,98,37,10,135,0,bf16 f16 f32 fp16 fp32 fp4_e2m1 fp6_e2m3 fp8_e4m3 nvfp4,directly_tested,135 direct test references
softmax_safe,4,4,1,8,0,bf16,directly_tested,8 direct test references
softplus,2,0,0,2,0,,directly_tested,2 direct test references
softplus,3,0,0,3,0,bf16 f32,directly_tested,3 direct test references
sort,3,0,0,3,0,,directly_tested,3 direct test references
spectral_conv,5,0,0,5,0,fp32,directly_tested,5 direct test references
spectral_filter,3,0,0,3,0,fp32,directly_tested,3 direct test references
spectral_norm,2,0,0,2,0,,directly_tested,2 direct test references
split,1,0,0,1,1,,structural_only,unclassified — defaults to structural_only
spmm_coo,3,0,0,3,0,,directly_tested,3 direct test references
spmm_csr,2,0,0,2,0,,directly_tested,2 direct test references
sqrt,4,0,0,4,0,,directly_tested,4 direct test references
sqrt,5,0,0,5,0,bf16 f32,directly_tested,5 direct test references
squeeze,1,0,0,1,1,,structural_only,unclassified — defaults to structural_only
stablehlo_export,2,0,0,2,0,,directly_tested,2 direct test references
stack,2,0,0,2,0,,directly_tested,2 direct test references
Expand All @@ -437,7 +437,7 @@ svd,8,5,0,13,0,bf16 f16 f32 fp16 fp32,directly_tested,13 direct test references
switch,0,0,0,0,1,,structural_only,category default for 'control_flow'
take,1,0,0,1,1,,structural_only,unclassified — defaults to structural_only
tan,1,0,0,1,1,,covered_by_family,category default for 'elementwise'
tanh,4,0,1,4,0,bf16 f16 f32,directly_tested,4 direct test references
tanh,5,0,1,5,0,bf16 f16 f32,directly_tested,5 direct test references
tile,1,0,0,1,1,,structural_only,unclassified — defaults to structural_only
tile_view,2,0,0,2,0,,directly_tested,2 direct test references
tiny_attention_conformance,0,0,0,0,1,,structural_only,category default for 'conformance'
Expand Down
Loading
Loading