Skip to content

Fix MXFP8 testing synchronization issue - #324

Closed
jhjpark wants to merge 4 commits into
NVIDIA:developfrom
jhjpark:jopark/fix_mxfp8_testing
Closed

Fix MXFP8 testing synchronization issue#324
jhjpark wants to merge 4 commits into
NVIDIA:developfrom
jhjpark:jopark/fix_mxfp8_testing

Conversation

@jhjpark

@jhjpark jhjpark commented Jun 25, 2026

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • New Features

    • Added support for cumulative sequence-length inputs, ragged offset multipliers, and additional grouped GEMM row-scaling options.
    • Expanded Python and C++ graph APIs with new plan/engine inspection and more flexible SDPA/reduction inputs.
  • Bug Fixes

    • Improved handling of dynamic shapes, stream selection, and version gating across multiple paths.
    • Fixed validation and serialization so more tensor and attention configurations work reliably.
  • Documentation

    • Updated README and benchmark docs with new examples, compatibility notes, and troubleshooting guidance.

Anerudhan and others added 4 commits June 8, 2026 12:03
…NVIDIA#277)

* bench: add autoregressive video DiT SDPA config + GB200/GB300 results

Adds a new benchmark config for the autoregressive (world-model / next-frame)
video DiT shape: short query (one new frame, s_q ∈ {985, 1024, 2048, 4096,
8192}) attending a long cached KV history (s_kv=62208) with h=9, d=128 and
no operator-level mask. This is a class of workload that prior DiT configs
(LTX-2, Wan 2.2) don't cover, because those run bidirectional self-attention
with s_q == s_kv.

Captured on lyris GB200 and GB300 (cuDNN 9.23.0, FAv4 from the CuTe-DSL
build). FAv4 FP8/MXFP8 bars are absent because that build's forward
asserts on non-fp16/bf16 inputs; the runner now skips FAv4 cases for both
FP8 and MXFP8 (previously only MXFP8) to keep the CSVs free of traceback
noise.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* bench: add B300 peak comparison for autoregressive DiT (cuDNN split-K vs FAv4 best num_splits)

Adds a "peak vs peak" view that complements the existing default-vs-default
chart: cuDNN 9.30.0 with prefill split-K enabled on bf16/fp8/mxfp8, paired
against FAv4 BF16 swept over num_splits ∈ {1, 2, 4, 8, 16, 32} with the
best per-seqlen result annotated on the bar (ks=).

For the autoregressive video DiT shape (B=1, h=9, d=128, s_q ∈ {985..8192},
s_kv=62208) on B300 SXM6:

  s_q   cuDNN BF16   cuDNN FP8   cuDNN MXFP8   FAv4 BF16 (best ks)
   985    1701          2429        2274         1424 (ks=4)
  1024    1767          2526        2367         1485 (ks=4)
  2048    1880          2713        2547         1597 (ks=2)
  4096    1997          2947        2655         1995 (ks=1)
  8192    1998          2974        2681         1980 (ks=1)
  (TFLOPS, fwd only)

cuDNN BF16+split-K beats FAv4-best-num_splits at every seqlen (+19% at the
short-Q end, tied at large s_q where neither needs splitting). FP8/MXFP8
dominate by +30-50% over FAv4 BF16 thanks to the higher mma throughput.

Changes:
  * benchmark_single_sdpa.py: --fa4_num_splits flag plumbed end-to-end so
    callers can force FAv4 into a specific split count (default unchanged:
    let FAv4 pick automatically).
  * bench_ar_dit_peak.py: standalone driver that runs the cartesian
    {seqlens} x {cudnn dtypes} sweep plus the FAv4 num_splits sweep and
    emits a CSV with one row per (backend, dtype, seqlen) — with the
    winning num_splits recorded for the FAv4 rows.
  * results/auto_regressive_dit/b300/: CSV + chart.
  * README: B300 peak section.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* bench: GB200 + GB300 peak comparison for autoregressive DiT (replace B300 preview)

Drops the earlier B300 preview chart in favour of the matching peak charts
on the production GB200 and GB300 superchip variants (same SM_103 silicon
in the GB300 case, fewer SMs / lower clock on GB200). Charts are the same
peak-vs-peak view: cuDNN 9.30.0 with prefill split-K enabled on
bf16/fp8/mxfp8, paired against FAv4 BF16 swept over num_splits and
keeping the best per-seqlen result.

GB300 (TFLOPS, fwd only):

  s_q   cuDNN BF16   cuDNN FP8   cuDNN MXFP8   FAv4 BF16 (best ks)
   985    1752          2519        2359          1451 (ks=4)
  1024    1813          2619        2447          1515 (ks=4)
  2048    1923          2768        2598          1613 (ks=2)
  4096    2050          2978        2687          2055 (ks=1)
  8192    2085          3002        2707          2071 (ks=1)

GB200 (TFLOPS, fwd only):

  s_q   cuDNN BF16   cuDNN FP8   cuDNN MXFP8   FAv4 BF16 (best ks)
   985    1380          1796        1717          1332 (ks=4)
  1024    1429          1870        1785          1389 (ks=4)
  2048    1573          1996        1915          1513 (ks=2)
  4096    1697          2066        1971          1746 (ks=1)
  8192    1762          2080        1988          1802 (ks=1)

On GB300 cuDNN BF16+split-K beats FAv4-best-num_splits at every seqlen
(+21% at the short-Q end, tied at large s_q where neither needs splitting).
On GB200 the short-Q advantage is +4-5% and FAv4 narrowly edges cuDNN BF16
at the large s_q end (-2-3%). FP8/MXFP8 dominate by +30-50% over FAv4
BF16 on both GPUs.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* bench: consolidate autoregressive DiT charts to a single canonical view per GPU

Drops the cuDNN 9.23 default-vs-default chart pair — those numbers are
stale relative to what ships next, and keeping two charts per GPU with
two different cuDNN versions is more confusing than informative. The
remaining chart on each GPU is the cuDNN 9.30.0 + prefill split-K view
paired against FAv4 BF16 with the best num_splits per seqlen, captured
on the production GB200 and GB300 superchips. CSV is named
auto_regressive_dit_no_mask.csv so the chart and its source data follow
the standard <config>_<mask>.{png,csv} convention used by other
benchmarks in this suite.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* bench: relabel autoregressive DiT charts to cuDNN 9.24.0 (split-K release version)

The split-K prefill feature exercised by these charts is cherry-picked
onto release/9.24.0 and ships in that release, so the chart labels and
the cudnn_backend_version column in the CSVs should reflect that
version rather than the dev-branch version they happened to be
measured on.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test/python: cap peak GPU memory via PYTORCH_CUDA_ALLOC_CONF (NVIDIA#247)

Long pytest-xdist runs (e.g. test_mhas_v2 ~2.5k SDPA configs in one
worker) hit a much higher GPU memory high-water mark than any single
test needs, because the caching allocator retains freed blocks across
configs.

Setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True,
garbage_collection_threshold:0.6 before torch is imported reduces the
peak to roughly the maximum any single test needs, with no change in
wall time or test outcome.

Use os.environ.setdefault so user-provided values still win, and
place it above the transformer_engine import so the env var is
visible by the time torch initializes its CUDA allocator.

* Fix DSA link in README.md

Updated the link for DSA in the README to point to the correct directory.

* Remove stale H200 benchmark artifacts (NVIDIA#252)

These artifacts were superseded by the newer SDPA benchmark result layout and were already removed from the internal GitLab develop branch.

* Change profile_pass from 'fwd' to 'both'

* Bump the develop to 1.25.0

* Fix varpack-template lifecycle bugs + add defensive checks

Two pre-existing bugs in the VariantPackTemplate, plus one defensive guard:

1. Graph copy -> dangling host pointers. template_ptrs stores raw addresses
   into cached_pass_by_value storage owned by the source Graph. Default copy
   propagated prepared=true while the addresses still pointed at the source.
   Fix: VarpackPrepStateBox copy ctor/assign now always start with
   prepared=false so the copy re-preps on first use against its own storage.
2. Re-deserialize on the same Graph -> stale template. deserialize(handle,...)
   rebinds cached_pass_by_value but the existing prepared=true causes the
   eager prep to short-circuit, leaving the slot layout from the prior
   deserialize. Fix: reset prepared=false and clear varpack_template before
   the eager prep call.
3. Null device_ptrs in raw-ptr create_variant_pack overloads. Reject nullptr
   + non-empty uids instead of forwarding to the cuDNN backend.

Adds explicit null-plan guards across detail::execute overloads, returning
GRAPH_EXECUTION_FAILED with "No plan found to execute!" instead of
dereferencing plan via plan->getTag().

Ports https://gitlab-master.nvidia.com/cudnn/cudnn_frontend/-/merge_requests/2117

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Clear deserialize-owned containers on re-deserialize

Addresses review feedback on PR NVIDIA#248: the prior fix reset prepared=false
and varpack_template but left deserialized_tensor_properties,
deserialized_pass_by_value, deserialized_workspace_modifications, and
tensors_to_dump populated from any earlier deserialize(handle, old_data).
On re-deserialize, prepare_variant_pack_template() could then ingest the
stale entries alongside the new ones.

Clear all four containers immediately after json::from_ubjson, before any
of the deserialize logic that repopulates them.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Add row-scale support to grouped GEMM quant

Signed-off-by: Ziang Li <ziangli@umich.edu>

* Tighten row-scale grouped GEMM quant tests

Signed-off-by: Ziang Li <ziangli@umich.edu>

* feat(python): add get_engine_and_knobs_at_index for structured plan pinning (NVIDIA#259)

* feat(python): add get_engine_and_knobs_at_index for structured plan pinning

get_plan_name_at_index returns a formatted "engN_kT=V" tag built from the
engine global index and knob choices. Callers that want to persist a tuned
plan and replay it later are forced to either store the bare plan index
(which drifts when the policy=ALL plan list is re-enumerated across
cudnn-frontend / backend versions) or parse the tag string.

Expose the structured data directly: get_engine_and_knobs_at_index returns
(engine_id, {KnobType_t: value}), reading the same backend attributes
get_engine_tag stringifies. The result feeds straight into
create_execution_plan(engine_id, knobs) to rebuild the exact same kernel on a
fresh graph without a heuristics query.

- detail::get_engine_id_and_knobs (cudnn_frontend_utils.h): structured reader
- Execution_plan_list::get_engine_and_knobs_at_index (plans.h)
- Graph::get_engine_and_knobs_at_index (graph_interface.h)
- PyGraph binding (pygraph.h/.cpp)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* address review: bounds-check index, add cpp unit test, trim comments

- get_engine_and_knobs_at_index: reject out-of-range index (mirrors
  check_support_at_index) instead of indexing engine_configs OOB.
- add test/cpp/get_engine_and_knobs.cpp: enumerate a matmul graph's plans,
  read (engine_id, knobs) for each, and confirm re-pinning via
  create_execution_plan reproduces the same plan (matching name); also checks
  out-of-range indices error.
- trim the new doc comments to match neighboring style.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* knobs: add SWAP_AB / INPUT_TMA_ENABLE / OUTPUT_TMA_ENABLE to KnobType_t

KnobType_t (and the to/from backend converters) stopped at WARP_SPEC_CFG (42),
so engines using SWAP_AB (43, cuDNN 9.18), INPUT_TMA_ENABLE (44) or
OUTPUT_TMA_ENABLE (45, cuDNN 9.22) had those knobs mapped to NOT_SET by
convert_from_backend_knob_type. Feeding NOT_SET back into create_execution_plan
then failed convert_to_backend_knob_type with INVALID_VALUE -- so a plan
enumerated with one of these knobs (e.g. via get_engine_and_knobs_at_index)
could not be pinned.

Add the three knob types to the enum, both converters (version-gated to match
the backend @SInCE), and the pybind knob_type enum.

The cpp test now compares the structured identity (engine id + knob map)
instead of the plan-name tag, since the tag serializes knobs in engine-config
order, which differs between the heuristic config and the pinned one even
though the kernel is identical. create_execution_plan is now asserted to
succeed for every enumerated plan; building it stays best-effort (can fail for
unrelated environment reasons such as a ptxas older than the engine's target).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* make get_engine_tag deterministic: sort knob choices by type

The plan-name tag was built by iterating CUDNN_ATTR_ENGINECFG_KNOB_CHOICES in
stored order, which differs between the heuristics path and
create_execution_plan (set_knob_choices iterates a std::unordered_map). So the
same engine + knob values could serialize to differently-ordered tags
(e.g. eng11_k2=29_k27=0...k43=0 vs eng11_k43=0_k38=0...k2=29) -- the kernel is
identical but the string isn't a stable id.

Sort the knob choices by type before formatting so the tag is a deterministic
function of the engine config regardless of how it was built. This is off the
execution hot path (tag is used for logging / plan identity), so no perf
impact; the actual knob choices passed to the backend are unchanged.

The cpp test now also asserts the pinned plan's tag matches the original's.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Yang Xu <yanxu@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Update SDPA Benchmarking Artifacts (NVIDIA#265)

* update sdpa benchmark artifacts

* update acknowledgement

* Adding coderabbit review guide (initial template)

* fix: allow overriding libcudart selection via CUDNN_FRONTEND_CUDART_LIB_NAME

When dynamic loading is enabled, load_cudart_so() searches for the supported
libcudart major versions and aborts with "Multiple libcudart libraries found"
when more than one is visible on the library search path. This happens in
containerized environments such as GKE, where the TCPXO NCCL plugin mounts a
different libcudart major version from the host than the one shipped in the
container.

Check the CUDNN_FRONTEND_CUDART_LIB_NAME environment variable first; when set
to a library name or path, dlopen exactly that library and skip the automatic
multi-version detection. Behavior is unchanged when the variable is unset.

Fixes NVIDIA#267

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Clean up guardword-flagged comments (xmma path, gitlab URL, P4 label, Perfsim, HACK/Ugly, STS/CGA SASS terms) (NVIDIA#273)

Comment-only cleanups, no behaviour change. Replaces guardword-flagged
phrasing with neutral equivalents in 7 files:

- attention_utils.h:67 — drop internal `xmma/fast_math.h:118-125` path
  reference; keep the rationale ("matches cuDNN backend's find_divisor_v2
  fast-math helper").
- test_sdpa_bwd.py:8 — drop `gitlab-master.nvidia.com` job URL from the
  module docstring; the rationale (2-CTA + Blackwell TMEM + xdist) is
  fully self-explanatory above it.
- dense_score_recompute_sm90.py — "Perfsim" → "Profiling";
  "Weights/LSE LDG" → "Weights/LSE load-from-global" (x2).
- indexer_backward_sm90.py — `# P4:` block-pass label → `# Pass 4:` (x2);
  rephrase 5 "STS" SASS-instruction references in comments to
  "shared-mem store(s)" / "write to shared mem".
- indexer_backward_sm100.py — same STS → shared-mem-store rephrasing
  in 1 docstring.
- dsa_bwd_sm90.py:386 — `# HACK:` → `# Note:` (same meaning).
- dsa_bwd_sm90.py:1554 — `STS(dS)` → "storing dS to shared mem".
- dsa_bwd_sm100.py:941 — `# Ugly,` → `# Awkward,`.
- dense_gemm_persistent_swiglu.py:1049 — "single CGA" → "single cluster".

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* remove_9.99_version_tag

* add_protection_flags

* fix(windows): consolidate getenv access and fix C4996/C4005 on MSVC

The Windows wheel build (deploy:build_bdist_wheels_3.10) failed because the
std::getenv call added to load_cudart_so() in cudnn_frontend_shim.h triggers
MSVC warning C4996 ('getenv' is unsafe), which is treated as an error under /WX.

Root cause and fixes:
- Move get_environment() to cudnn_frontend_shim.h (the lowest-level header,
  included by utils.h before Logging.h) so a single definition is shared by all
  layers without inverting include dependencies. It wraps std::getenv with a
  properly scoped #pragma warning(push)/disable(4996)/pop, guarded by _WIN32.
- Route all getenv call sites through get_environment(): shim.h, graph_properties.h,
  scaled_dot_product_flash_attention.h, and sm100_rms_norm_silu_engine.h. These were
  previously only spared from C4996 by an unscoped pragma leak in Logging.h, and would
  have started failing once that leak was fixed.
- Remove the duplicate get_environment() from cudnn_frontend_Logging.h, which had three
  issues: an unscoped 'warning(disable:4996)' that leaked to the rest of the TU, a
  no-op '#define _CRT_SECURE_NO_WARNINGS' (placed after the CRT headers), and a 'WIN32'
  guard that should be '_WIN32'. Dropping the macro also resolves the C4005
  '_CRT_SECURE_NO_WARNINGS macro redefinition' warning for downstream projects.

Fixes NVIDIA#139

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(shim): warn instead of throwing when multiple libcudart libraries are found

Loading cudart no longer aborts when both libcudart.so.12 and libcudart.so.13
are present in the library search path. Instead, load_cudart_so() emits a
warning on stderr and falls back to the first library found. Users can still
select a specific library explicitly via CUDNN_FRONTEND_CUDART_LIB_NAME.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Unblock SDPA tests and promote FP8 ragged backward to L0 (NVIDIA#275)

* Promote L1 Python tests to L0

* Restore L1 markers except FP8 ragged backward

* Add per-expert reduction (group_offset) for MoE grouped GEMM

Adds optional group_offset support to the reduction node so cuDNN FE can
express per-expert reductions for MoE grouped GEMM workloads.

- New Group_offset graph_properties tensor input and
  Reduction_attributes::set_group_offset setter
- INode::reduction and PyGraph::reduction signatures take an optional
  group_offset tensor
- Operation_v8 builder wires CUDNN_ATTR_OPERATION_REDUCTION_GROUP_OFFSET_DESC
  with runtime version checks (cuDNN >= 9.24.0)
- Python binding (pygraph) exposes the optional group_offset argument

Mirrors gitlab-master cudnn/cudnn_frontend MR !2111 by @yanqinz.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Fix the 9.99 bound

* Skip flexible-graph SDPA bwd sample on SM120 and above (NVIDIA#284)

The fp16 backward-with-flexible-graphs sample guards against SM 120
(consumer Blackwell) where this path is not supported. The guard used
an exact == 120 check, which missed SM 121 (GB10 / DGX Spark) and any
later consumer Blackwell arch, causing the sample to run and fail there.

Change the check to >= 120 so the sample is skipped on SM 120 and above,
and update the SKIP message to match.

Co-authored-by: Yang Xu <yanxu@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* 1

* Add pre-commit hooks (NVIDIA#286)

* Fix clang format issues

* Fix clang-format

* Add pre-commit hooks and fix pre-commit

* Fix the black issues

* Skip TensorIR MemBound / compile-time-const samples on consumer Blackwell (SM12x) (NVIDIA#285)

* Skip TensorIR MemBound / compile-time-const samples on consumer Blackwell (SM12x)

The TensorIR MemBound engine (cudnnTensorIrMemBoundEngine) only supports
SM100-SM109 (data center Blackwell): its arch gate is [SM_100, SM_110) and the
DKG cubins it emits are the sm_100f family-portable target, which the CUDA
driver will not load on sm_120. The membound and compile-time-constant samples
guarded their device check with check_device_arch_newer_than("blackwell") /
is_blackwell_arch(), both of which are true for SM120 consumer Blackwell. So on
an RTX 50-series (sm_120) GPU these samples fall through to
create_execution_plans() and FAIL with "No valid engine configs returned from
heuristics" (no engine serves the graph; the kernelgen runtime-fusion fallback
only targets SM70/SM80/SM90).

Narrow the guard to is_blackwell_computing_arch() (100 <= cc < 110) so the
samples skip cleanly on SM120 and above, matching the backend engine's actual
support range. This mirrors PR NVIDIA#283, which skipped the flexible-graph SDPA
backward sample on SM120+.

Affected test cases (verified on RTX 5080 / sm_120, cuDNN 9.30 -> now SKIP):
  membound/transpose.cpp        "Membound transpose permutes dims"
  membound/reshape.cpp          "Membound reshape ... LOGICAL mode"
  membound/slice.cpp            "Membound slice window with step"
  membound/concat.cpp           "Membound concatenate on channel axis"
  membound/membound_fusion.cpp  "Fusion reshape then ReLU" / "Fusion transpose then add bias tensor"
  membound/boolean_fusion.cpp   "Boolean CMP_GT and LOGICAL_AND fusion"
  misc/compile_time_constant_example.cpp  "Compile-time constant scalar multiply and add"

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Skip boolean_cmp_logic Python notebook on consumer Blackwell (SM12x)

Python counterpart of the C++ membound/boolean sample fix. The CMP_GT +
LOGICAL_AND boolean fusion runs on the TensorIR mem-bound engine, which only
supports SM100-SM109 (data center Blackwell). On SM120 consumer Blackwell the
notebook's create_execution_plans([A, FALLBACK]) silently falls back to an
engine that produces WRONG results (verified on RTX 5080 / sm_120: 109/512
mismatches -> assertion failure).

Gate the cuDNN cells on is_supported_arch so the notebook skips cleanly on
SM120 instead of producing wrong results, and fix the prerequisite markdown
(SM100+ "or later" -> SM100-SM109). The arch check computes the full compute
capability (major*10 + minor) and tests 100 <= cc < 110 to mirror the C++
is_blackwell_computing_arch() helper exactly.

This notebook is not part of ci/run_python_samples.sh, so it does not affect
CI; the fix is for correctness/consistency with the C++ sample.

Committed with --no-verify: the local black-jupyter pre-commit hook reflows the
whole .ipynb to indent=1 (repo notebooks are indent=2) and collapses unrelated
aligned dicts; CI does not enforce notebook formatting.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Yang Xu <yanxu@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Support cu_seqlens in unified SDPA (NVIDIA#266)

* use static signature for sfd_col_d_srelu_tensor (NVIDIA#281)

Signed-off-by: Jieming Zhang <jiemingz@nvidia.com>

* DSA: fix CuTe DSL guards and add SM90 indexer forward (NVIDIA#263)

* DSA: fix CuTe DSL guards and add SM90 indexer forward

* DSA: allow indexer top-k on SM90

* DSA: trim CuTe DSL compile-cache keys + unify indexer_forward paths

Compile-cache keys across the deepseek_sparse_attention kernels included
runtime-only values (batch/seqlen/seqlen_k, sm_scale, tensor shapes/strides,
num_head, num_threads), forcing spurious recompiles under varlen / changing
batch even though one compiled kernel serves them all. Drop those fields and
keep only params that change generated code.

The two dense_indexer_backward kernels originally baked seqlen into codegen,
so to drop it safely they were reworked to take seqlen at runtime:
  - sm90: the dense K-load looped via range_constexpr(num_topk_blocks =
    seqlen_k // block_I); it now loops at runtime over num_k_blocks, like the
    compute warpgroup already did.
  - sm100: ScoreGradDense baked max_seqlen_q into its launch grid and
    max_seqlen_q/k into the causal-mask bound via __init__ ints; they are now
    runtime Int32 args (matching the GEMM kernel), which also fixes a latent
    bug where a kernel compiled for one max_seqlen_k could be silently reused
    for another.

Collapse the redundant two-layer compile cache (dict-of-closures + per-closure
lazy holder) in the indexer_backward factories to the single forward-style dict
(key -> compiled kernel), matching indexer_forward.

indexer_forward: route the SM100 BSHD path through the same indexer_fwd wrapper
as THD instead of the separate IndexerForward APIBase class, which compiled
against concrete fake-tensor shapes (recompiling per shape/stride). indexer_fwd
marks layouts dynamic and compiles once per config; on B300 the two produce
bit-identical output with <2% kernel-time difference at realistic shapes.
indexer_fwd gains an optional current_stream arg (also fixing the THD path,
which previously dropped the caller's stream). The public IndexerForward
class/export is retained.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* DSA: address indexer stream and cache review

* DSA: format CuTe DSL indexer files

* DSA: key SM100 sparse bwd by num heads

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: mingyangw <mingyangw@nvidia.com>

* Fix formatting issues from NVIDIA#263 (NVIDIA#294)

* Support static linking of libcudnn (NVIDIA#182)

* Support static linking of libcudnn

* Fix variable handling

* Don't use static zlib for PIC

* Rename CUDNN_STATIC_LINK

* Make version variables compatible for pytorch

* Apply suggestion from @coderabbitai[bot]

Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>

* Apply review suggestions

---------

Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>

* make dgeglu config values compile time constants instead of runtime values (NVIDIA#293)

* bench: add autoregressive video DiT SDPA config + GB200/GB300 results (NVIDIA#277) (NVIDIA#295)

* bench: add autoregressive video DiT SDPA config + GB200/GB300 results

Adds a new benchmark config for the autoregressive (world-model / next-frame)
video DiT shape: short query (one new frame, s_q ∈ {985, 1024, 2048, 4096,
8192}) attending a long cached KV history (s_kv=62208) with h=9, d=128 and
no operator-level mask. This is a class of workload that prior DiT configs
(LTX-2, Wan 2.2) don't cover, because those run bidirectional self-attention
with s_q == s_kv.

Captured on lyris GB200 and GB300 (cuDNN 9.23.0, FAv4 from the CuTe-DSL
build). FAv4 FP8/MXFP8 bars are absent because that build's forward
asserts on non-fp16/bf16 inputs; the runner now skips FAv4 cases for both
FP8 and MXFP8 (previously only MXFP8) to keep the CSVs free of traceback
noise.



* bench: add B300 peak comparison for autoregressive DiT (cuDNN split-K vs FAv4 best num_splits)

Adds a "peak vs peak" view that complements the existing default-vs-default
chart: cuDNN 9.30.0 with prefill split-K enabled on bf16/fp8/mxfp8, paired
against FAv4 BF16 swept over num_splits ∈ {1, 2, 4, 8, 16, 32} with the
best per-seqlen result annotated on the bar (ks=).

For the autoregressive video DiT shape (B=1, h=9, d=128, s_q ∈ {985..8192},
s_kv=62208) on B300 SXM6:

  s_q   cuDNN BF16   cuDNN FP8   cuDNN MXFP8   FAv4 BF16 (best ks)
   985    1701          2429        2274         1424 (ks=4)
  1024    1767          2526        2367         1485 (ks=4)
  2048    1880          2713        2547         1597 (ks=2)
  4096    1997          2947        2655         1995 (ks=1)
  8192    1998          2974        2681         1980 (ks=1)
  (TFLOPS, fwd only)

cuDNN BF16+split-K beats FAv4-best-num_splits at every seqlen (+19% at the
short-Q end, tied at large s_q where neither needs splitting). FP8/MXFP8
dominate by +30-50% over FAv4 BF16 thanks to the higher mma throughput.

Changes:
  * benchmark_single_sdpa.py: --fa4_num_splits flag plumbed end-to-end so
    callers can force FAv4 into a specific split count (default unchanged:
    let FAv4 pick automatically).
  * bench_ar_dit_peak.py: standalone driver that runs the cartesian
    {seqlens} x {cudnn dtypes} sweep plus the FAv4 num_splits sweep and
    emits a CSV with one row per (backend, dtype, seqlen) — with the
    winning num_splits recorded for the FAv4 rows.
  * results/auto_regressive_dit/b300/: CSV + chart.
  * README: B300 peak section.



* bench: GB200 + GB300 peak comparison for autoregressive DiT (replace B300 preview)

Drops the earlier B300 preview chart in favour of the matching peak charts
on the production GB200 and GB300 superchip variants (same SM_103 silicon
in the GB300 case, fewer SMs / lower clock on GB200). Charts are the same
peak-vs-peak view: cuDNN 9.30.0 with prefill split-K enabled on
bf16/fp8/mxfp8, paired against FAv4 BF16 swept over num_splits and
keeping the best per-seqlen result.

GB300 (TFLOPS, fwd only):

  s_q   cuDNN BF16   cuDNN FP8   cuDNN MXFP8   FAv4 BF16 (best ks)
   985    1752          2519        2359          1451 (ks=4)
  1024    1813          2619        2447          1515 (ks=4)
  2048    1923          2768        2598          1613 (ks=2)
  4096    2050          2978        2687          2055 (ks=1)
  8192    2085          3002        2707          2071 (ks=1)

GB200 (TFLOPS, fwd only):

  s_q   cuDNN BF16   cuDNN FP8   cuDNN MXFP8   FAv4 BF16 (best ks)
   985    1380          1796        1717          1332 (ks=4)
  1024    1429          1870        1785          1389 (ks=4)
  2048    1573          1996        1915          1513 (ks=2)
  4096    1697          2066        1971          1746 (ks=1)
  8192    1762          2080        1988          1802 (ks=1)

On GB300 cuDNN BF16+split-K beats FAv4-best-num_splits at every seqlen
(+21% at the short-Q end, tied at large s_q where neither needs splitting).
On GB200 the short-Q advantage is +4-5% and FAv4 narrowly edges cuDNN BF16
at the large s_q end (-2-3%). FP8/MXFP8 dominate by +30-50% over FAv4
BF16 on both GPUs.



* bench: consolidate autoregressive DiT charts to a single canonical view per GPU

Drops the cuDNN 9.23 default-vs-default chart pair — those numbers are
stale relative to what ships next, and keeping two charts per GPU with
two different cuDNN versions is more confusing than informative. The
remaining chart on each GPU is the cuDNN 9.30.0 + prefill split-K view
paired against FAv4 BF16 with the best num_splits per seqlen, captured
on the production GB200 and GB300 superchips. CSV is named
auto_regressive_dit_no_mask.csv so the chart and its source data follow
the standard <config>_<mask>.{png,csv} convention used by other
benchmarks in this suite.



* bench: relabel autoregressive DiT charts to cuDNN 9.24.0 (split-K release version)

The split-K prefill feature exercised by these charts is cherry-picked
onto release/9.24.0 and ships in that release, so the chart labels and
the cudnn_backend_version column in the CSVs should reflect that
version rather than the dev-branch version they happened to be
measured on.



---------

Co-authored-by: Vedaanta Agarwalla <142048820+vedaanta@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* - Update the Black version. (NVIDIA#296)

- Fix the formatting issues in grouped_gemm_dglu/api.py

* Add ragged offset multiplier support (NVIDIA#290)

Add frontend support for the per-tensor ragged offset multiplier
(CUDNN_ATTR_TENSOR_RAGGED_OFFSET_MULTIPLIER), letting ragged offsets be
stored in coarser units and scaled back to element offsets by the engine.

- Add ragged_offset_multiplier field, getters/setters, and validation to
  Tensor_attributes; emit the backend attribute (gated on cuDNN >= 9.24.0).
- Expose ragged_offset_multiplier through the Python tensor() bindings
  (appended last to preserve positional backward compatibility).
- Serialize/deserialize the multiplier and the ragged offset reference.
- Reject a non-default multiplier on the composite SDPA path (unified
  forward only).
- Add C++ and Python (test_mhas_v2) coverage, including a cu_ragged_mult
  configuration exercising cu_seqlens together with the multiplier.

* Fix unused ragged offset version error variable (NVIDIA#299)

`NV_CUDNN_FE_DYNAMIC_CHECK_BACKEND_DESCRIPTOR` expands to nothing when
`NV_CUDNN_FRONTEND_USE_DYNAMIC_LOADING` is not defined. So, the variable
`ragged_offset_multiplier_cudnn_ver_error` may be unused.

---------

Signed-off-by: Ziang Li <ziangli@umich.edu>
Signed-off-by: Jieming Zhang <jiemingz@nvidia.com>
Co-authored-by: Vedaanta Agarwalla <142048820+vedaanta@users.noreply.github.com>
Co-authored-by: Hwanseo Choi <hwanseoc@nvidia.com>
Co-authored-by: Vincent <vinnietombari@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Ziang Li <ziangli@umich.edu>
Co-authored-by: Yang Xu <38851819+YangXu1990uiuc@users.noreply.github.com>
Co-authored-by: Yang Xu <yanxu@nvidia.com>
Co-authored-by: Brandon Zhang <31413216+brandonfzhang@users.noreply.github.com>
Co-authored-by: Jane (Jiancheng) Liu <liujane@nvidia.com>
Co-authored-by: Yanqin Zhai <yanqinz@nvidia.com>
Co-authored-by: Emil Gilliam <egilliam@nvidia.com>
Co-authored-by: Jimmy Zhang <133159885+jiemingz@users.noreply.github.com>
Co-authored-by: jiayus-nvidia <jiayus@nvidia.com>
Co-authored-by: mingyangw <mingyangw@nvidia.com>
Co-authored-by: Takeshi Watanabe <take-cheeze@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: Mingyang Wang <35635157+saltyminty@users.noreply.github.com>
Co-authored-by: Shraiysh <svaishay@nvidia.com>
@coderabbitai

coderabbitai Bot commented Jun 25, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 442d2251-29d1-41aa-8e0c-aede70118e1e

📥 Commits

Reviewing files that changed from the base of the PR and between c4ec01a and 4606002.

⛔ Files ignored due to path filters (79)
  • benchmark/sdpa_benchmark_training/results/auto_regressive_dit/gb200/auto_regressive_dit_no_mask.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/auto_regressive_dit/gb200/auto_regressive_dit_no_mask.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/auto_regressive_dit/gb300/auto_regressive_dit_no_mask.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/auto_regressive_dit/gb300/auto_regressive_dit_no_mask.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/dsv3/gb200/dsv3_20260424_101009.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/dsv3/gb200/dsv3_20260529_181100.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/dsv3/gb200/dsv3_no_mask.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/dsv3/gb200/dsv3_no_mask_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/dsv3/gb200/dsv3_top_left.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/dsv3/gb200/dsv3_top_left_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/dsv3/gb300/dsv3_20260424_101002.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/dsv3/gb300/dsv3_20260529_175553.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/dsv3/gb300/dsv3_no_mask.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/dsv3/gb300/dsv3_no_mask_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/dsv3/gb300/dsv3_top_left.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/dsv3/gb300/dsv3_top_left_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/gpt_oss/gb200/gpt_oss_20260424_100011.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/gpt_oss/gb200/gpt_oss_20260529_180050.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/gpt_oss/gb200/gpt_oss_top_left.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/gpt_oss/gb200/gpt_oss_top_left_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/gpt_oss/gb300/gpt_oss_20260424_100022.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/gpt_oss/gb300/gpt_oss_20260529_174551.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/gpt_oss/gb300/gpt_oss_top_left.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/gpt_oss/gb300/gpt_oss_top_left_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/h200_919_only_cudnn/dsv3_20260227_034744.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/h200_919_only_cudnn/dsv3_top_left.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/h200_919_only_cudnn/gpt_oss_20260227_034819.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/h200_919_only_cudnn/gpt_oss_top_left.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/h200_919_only_cudnn/llama3.1_20260227_034703.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/h200_919_only_cudnn/llama3.1_no_mask.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/h200_919_only_cudnn/llama3.1_top_left.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/kimiK26/gb200/kimiK26_20260424_100953.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/kimiK26/gb200/kimiK26_20260529_181016.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/kimiK26/gb200/kimiK26_no_mask.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/kimiK26/gb200/kimiK26_no_mask_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/kimiK26/gb200/kimiK26_top_left.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/kimiK26/gb200/kimiK26_top_left_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/kimiK26/gb300/kimiK26_20260424_100915.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/kimiK26/gb300/kimiK26_20260529_175511.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/kimiK26/gb300/kimiK26_no_mask.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/kimiK26/gb300/kimiK26_no_mask_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/kimiK26/gb300/kimiK26_top_left.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/kimiK26/gb300/kimiK26_top_left_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/llama3.1/gb200/llama3.1_20260424_100750.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/llama3.1/gb200/llama3.1_20260529_180853.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/llama3.1/gb200/llama3.1_no_mask.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/llama3.1/gb200/llama3.1_no_mask_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/llama3.1/gb200/llama3.1_top_left.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/llama3.1/gb200/llama3.1_top_left_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/llama3.1/gb300/llama3.1_20260424_100757.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/llama3.1/gb300/llama3.1_20260529_175350.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/llama3.1/gb300/llama3.1_no_mask.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/llama3.1/gb300/llama3.1_no_mask_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/llama3.1/gb300/llama3.1_top_left.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/llama3.1/gb300/llama3.1_top_left_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/ltx2/gb200/ltx2_20260424_095758.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/ltx2/gb200/ltx2_20260529_181611.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/ltx2/gb200/ltx2_no_mask.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/ltx2/gb200/ltx2_no_mask_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/ltx2/gb300/ltx2_20260424_095719.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/ltx2/gb300/ltx2_20260529_180103.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/ltx2/gb300/ltx2_no_mask.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/ltx2/gb300/ltx2_no_mask_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/qwen35/gb200/qwen35_20260424_095249.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/qwen35/gb200/qwen35_20260529_180715.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/qwen35/gb200/qwen35_top_left.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/qwen35/gb200/qwen35_top_left_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/qwen35/gb300/qwen35_20260424_095247.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/qwen35/gb300/qwen35_20260529_175216.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/qwen35/gb300/qwen35_top_left.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/qwen35/gb300/qwen35_top_left_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/wan22/gb200/wan22_20260424_095743.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/wan22/gb200/wan22_20260529_181549.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/wan22/gb200/wan22_no_mask.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/wan22/gb200/wan22_no_mask_det_overhead.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/wan22/gb300/wan22_20260424_095741.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/wan22/gb300/wan22_20260529_180039.csv is excluded by !**/*.csv
  • benchmark/sdpa_benchmark_training/results/wan22/gb300/wan22_no_mask.png is excluded by !**/*.png
  • benchmark/sdpa_benchmark_training/results/wan22/gb300/wan22_no_mask_det_overhead.png is excluded by !**/*.png
📒 Files selected for processing (105)
  • .coderabbit.yaml
  • .pre-commit-config.yaml
  • CMakeLists.txt
  • README.md
  • benchmark/sdpa_benchmark_training/ACKNOWLEDGEMENTS.md
  • benchmark/sdpa_benchmark_training/README.md
  • benchmark/sdpa_benchmark_training/bench_ar_dit_peak.py
  • benchmark/sdpa_benchmark_training/benchmark_single_sdpa.py
  • benchmark/sdpa_benchmark_training/charts.py
  • benchmark/sdpa_benchmark_training/configs/auto_regressive_dit.py
  • benchmark/sdpa_benchmark_training/configs/qwen35.py
  • benchmark/sdpa_benchmark_training/runner.py
  • cmake/cuDNN.cmake
  • include/cudnn_frontend/backend/execution_helpers.h
  • include/cudnn_frontend/cudnn_interface.h
  • include/cudnn_frontend/experimental/attention_utils.h
  • include/cudnn_frontend/experimental/sm100_rms_norm_silu_engine.h
  • include/cudnn_frontend/graph_interface.h
  • include/cudnn_frontend/graph_properties.h
  • include/cudnn_frontend/knobs.h
  • include/cudnn_frontend/node/diagonal_band_mask.h
  • include/cudnn_frontend/node/moe_grouped_matmul_bwd.h
  • include/cudnn_frontend/node/reduction.h
  • include/cudnn_frontend/node/scaled_dot_product_flash_attention.h
  • include/cudnn_frontend/node/sdpa_fp8_bwd.h
  • include/cudnn_frontend/node/sdpa_support_surface.h
  • include/cudnn_frontend/node/softmax.h
  • include/cudnn_frontend/node_interface.h
  • include/cudnn_frontend/plans.h
  • include/cudnn_frontend/utils/attn_score_modifiers.h
  • include/cudnn_frontend/utils/serialize.h
  • include/cudnn_frontend_Logging.h
  • include/cudnn_frontend_Operation.h
  • include/cudnn_frontend_Tensor.h
  • include/cudnn_frontend_shim.h
  • include/cudnn_frontend_utils.h
  • include/cudnn_frontend_version.h
  • python/cudnn/__init__.py
  • python/cudnn/deepseek_sparse_attention/README.md
  • python/cudnn/deepseek_sparse_attention/indexer_backward/dense_indexer_backward_sm100.py
  • python/cudnn/deepseek_sparse_attention/indexer_backward/dense_indexer_backward_sm90.py
  • python/cudnn/deepseek_sparse_attention/indexer_backward/indexer_backward_sm100.py
  • python/cudnn/deepseek_sparse_attention/indexer_backward/indexer_backward_sm90.py
  • python/cudnn/deepseek_sparse_attention/indexer_forward/_interface.py
  • python/cudnn/deepseek_sparse_attention/indexer_forward/_interface_sm90.py
  • python/cudnn/deepseek_sparse_attention/indexer_forward/api.py
  • python/cudnn/deepseek_sparse_attention/indexer_forward/indexer_fwd_sm90.py
  • python/cudnn/deepseek_sparse_attention/indexer_top_k/api.py
  • python/cudnn/deepseek_sparse_attention/indexer_top_k/indexer_top_k_decode_varlen.py
  • python/cudnn/deepseek_sparse_attention/indexer_top_k/local_to_global_dsl.py
  • python/cudnn/deepseek_sparse_attention/score_recompute/_interface_sm100.py
  • python/cudnn/deepseek_sparse_attention/score_recompute/_interface_sm90.py
  • python/cudnn/deepseek_sparse_attention/score_recompute/dense_score_recompute_sm90.py
  • python/cudnn/deepseek_sparse_attention/score_recompute/pack_gqa.py
  • python/cudnn/deepseek_sparse_attention/score_recompute/sparse_score_recompute_sm100.py
  • python/cudnn/deepseek_sparse_attention/sparse_attention_backward/_interface_sm100.py
  • python/cudnn/deepseek_sparse_attention/sparse_attention_backward/_interface_sm90.py
  • python/cudnn/deepseek_sparse_attention/sparse_attention_backward/api.py
  • python/cudnn/deepseek_sparse_attention/sparse_attention_backward/dsa_bwd_sm100.py
  • python/cudnn/deepseek_sparse_attention/sparse_attention_backward/dsa_bwd_sm90.py
  • python/cudnn/gemm_swiglu/dense_gemm_persistent_swiglu.py
  • python/cudnn/grouped_gemm/grouped_gemm_dglu/api.py
  • python/cudnn/grouped_gemm/grouped_gemm_dglu/moe_blockscaled_grouped_gemm_dglu_dbias.py
  • python/cudnn/grouped_gemm/grouped_gemm_dsrelu/api.py
  • python/cudnn/grouped_gemm/grouped_gemm_quant/api.py
  • python/cudnn/grouped_gemm/grouped_gemm_quant/grouped_gemm_quant.py
  • python/cudnn/grouped_gemm/moe_sched_extension.py
  • python/properties.cpp
  • python/pygraph/pygraph.cpp
  • python/pygraph/pygraph.h
  • python/pygraph/sdpa.cpp
  • samples/cpp/CMakeLists.txt
  • samples/cpp/matmul/blackwell_nvfp4_mxfp8_block_scale_matmul.cpp
  • samples/cpp/matmul/matmuls.cpp
  • samples/cpp/membound/boolean_fusion.cpp
  • samples/cpp/membound/concat.cpp
  • samples/cpp/membound/membound_fusion.cpp
  • samples/cpp/membound/reshape.cpp
  • samples/cpp/membound/slice.cpp
  • samples/cpp/membound/transpose.cpp
  • samples/cpp/misc/compile_time_constant_example.cpp
  • samples/cpp/moe_grouped_matmul/moe_grouped_matmul.cpp
  • samples/cpp/sdpa/fp16_bwd_with_flexible_graphs.cpp
  • samples/cpp/sdpa/fp16_dynamic_shapes.cpp
  • samples/cpp/sdpa/fp16_fwd_with_cu_seq_len.cpp
  • samples/python/70_boolean_cmp_logic.ipynb
  • test/cpp/CMakeLists.txt
  • test/cpp/get_engine_and_knobs.cpp
  • test/cpp/tensor.cpp
  • test/python/conftest.py
  • test/python/fe_api/dsa/dsa_reference.py
  • test/python/fe_api/dsa/test_DSA_indexer_forward.py
  • test/python/fe_api/dsa/test_DSA_indexer_top_k.py
  • test/python/fe_api/test_grouped_gemm_quant.py
  • test/python/fe_api/test_grouped_gemm_quant_utils.py
  • test/python/fe_api/test_sdpa_bwd.py
  • test/python/sdpa/blocked.py
  • test/python/sdpa/fp16.py
  • test/python/sdpa/fp8.py
  • test/python/sdpa/mxfp8.py
  • test/python/sdpa/random_config.py
  • test/python/test_block_scale_quantize_dynamic_shape.py
  • test/python/test_matmul_bias_relu.py
  • test/python/test_mhas_v2.py
  • test/python/test_moe_grouped_matmul.py

📝 Walkthrough

Walkthrough

The PR adds frontend support for ragged-offset multipliers, group offsets, and cu-seq-len tensors, expands SDPA and DeepSeek sparse-attention paths, adds grouped GEMM row-scale and dGeGLU controls, and updates related benchmarks, samples, and tests.

Changes

Frontend runtime and build support

Layer / File(s) Summary
Build and loader support
*.yaml, CMakeLists.txt, README.md, cmake/cuDNN.cmake, include/cudnn_frontend/..., python/cudnn/__init__.py
Build metadata, version values, cuDNN discovery/loading, and runtime environment helpers are updated together.

Tensor ragged-offset metadata

Layer / File(s) Summary
Ragged offset multiplier
include/cudnn_frontend/graph_properties.h, include/cudnn_frontend/Tensor.h, include/cudnn_frontend/cudnn_interface.h, include/cudnn_frontend/utils/serialize.h, python/cudnn/__init__.py, python/properties.cpp, python/pygraph/..., test/cpp/tensor.cpp
Tensor attributes carry, serialize, bind, and validate ragged_offset_multiplier through the frontend and Python layers.

Reduction, knobs, and execution-plan queries

Layer / File(s) Summary
Reduction group offset
include/cudnn_frontend/graph_properties.h, include/cudnn_frontend/node_interface.h, include/cudnn_frontend/node/reduction.h, include/cudnn_frontend/Operation.h, python/pygraph/..., python/properties.cpp
Reduction builders accept an optional group_offset tensor and wire it into the backend descriptor.
Engine knobs and plan metadata
include/cudnn_frontend/graph_interface.h, include/cudnn_frontend/plans.h, include/cudnn_frontend/knobs.h, include/cudnn_frontend_utils.h, include/cudnn_frontend/backend/execution_helpers.h, python/pygraph/..., test/cpp/get_engine_and_knobs.cpp
Plan metadata now returns engine ids and knob maps, and the new knob types are bound through C++ and Python.

SDPA cu_seq_len surface and validation

Layer / File(s) Summary
cu_seq_len graph contracts
include/cudnn_frontend/..., python/pygraph/sdpa.cpp, python/pygraph/pygraph.h
SDPA graph attributes, node validation, serialization, and Python bindings accept cumulative sequence-length tensors.
Benchmarks and sample wiring
benchmark/sdpa_benchmark_training/*, samples/cpp/CMakeLists.txt, samples/cpp/sdpa/*
SDPA benchmark configs, sample entry points, and result documentation add cu_seq_len and autoregressive variants.
Tests and fixtures
test/python/sdpa/*, test/python/conftest.py
SDPA tests and fixtures update tensor shapes, capability gates, and blocked-case handling.

DeepSeek sparse attention

Layer / File(s) Summary
Forward indexer path
python/cudnn/deepseek_sparse_attention/indexer_forward/*, indexer_top_k/*, README.md
The forward indexer dispatches across SM90 and SM100 kernels and updates the SM90 top-k kernel and cache keys.
Backward and score recompute
python/cudnn/deepseek_sparse_attention/indexer_backward/*, score_recompute/*, sparse_attention_backward/*
Backward interfaces, dense kernels, and score recompute paths update stream handling, compile keys, and cached launches.
DSA tests
test/python/fe_api/dsa/*
The DSA reference and indexer tests update causal masks, capability gates, and supported exception handling.

Grouped GEMM scaling and MoE

Layer / File(s) Summary
dGeGLU and MoE grouped matmul
python/cudnn/grouped_gemm/grouped_gemm_dglu/*, python/cudnn/grouped_gemm/grouped_gemm_dsrelu/api.py, include/cudnn_frontend/node/moe_grouped_matmul_bwd.h, samples/cpp/moe_grouped_matmul/*, test/python/test_moe_grouped_matmul.py
dGeGLU parameters move into compile-time kernel construction, and MoE grouped matmul gating and docs update with the new limits.
Row-scale quant path
python/cudnn/grouped_gemm/grouped_gemm_quant/*, python/cudnn/grouped_gemm/moe_sched_extension.py, test/python/fe_api/test_grouped_gemm_quant*.py
Grouped GEMM quant kernels, scheduler extensions, and tests thread per-row scaling through compile, execute, and reference paths.

Sample and test architecture gating

Layer / File(s) Summary
Architecture gating
samples/cpp/matmul/*, samples/cpp/membound/*, samples/cpp/misc/compile_time_constant_example.cpp, samples/python/70_boolean_cmp_logic.ipynb, test/python/test_block_scale_quantize_dynamic_shape.py, test/python/test_matmul_bias_relu.py
Sample and test skip conditions move to Blackwell compute-architecture checks and broadened backend-version branches.

Estimated code review effort

🎯 5 (Critical) | ⏱️ ~90+ minutes

Possibly related PRs

  • NVIDIA/cudnn-frontend#263: Shares the DeepSeek sparse-attention kernel and cache-key areas, including indexer backward and forward-path changes.
  • NVIDIA/cudnn-frontend#317: Touches the same sparse score-recompute epilogue path and gating logic as the sparse_score_recompute_sm100.py updates.

Suggested labels

cat-enhancements

Suggested reviewers

  • saltyminty
  • Anerudhan

Poem

A rabbit hops through kernels bright,
With cu-seq-lens and row scales right.
I twitch my nose at plans and knobs,
Then nibble carrots, done with jobs. 🐇

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
⚔️ Resolve merge conflicts
  • Resolve merge conflict in branch jopark/fix_mxfp8_testing

Comment @coderabbitai help to get the list of available commands.

@jhjpark jhjpark closed this Jun 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants