Skip to content

fix(kernel): drop Triton vendored CUDA include path from nvcc -I - #247

Merged
lightseek-bot merged 1 commit into
mainfrom
zhyncs/kernel-drop-triton-include-path
May 25, 2026
Merged

lightseek-bot merged 1 commit into
mainfrom
zhyncs/kernel-drop-triton-include-path

Conversation

@zhyncs

@zhyncs zhyncs commented May 25, 2026 •

Copy link
Copy Markdown
Member

_resolve_include_dirs folds tokenspeed_triton/backends/nvidia/include and triton/backends/nvidia/include into nvcc's -I. Those bundles ship a CUDA-12-era crt/host_runtime.h whose __cudaLaunch is a 1-arg macro; nvcc 13.0's generated stubs call it with 2 args. On hosts where the system CUDA's crt/host_runtime.h isn't on nvcc's implicit search path (e.g. apt-installed CUDA 13 SDK behind a /usr/local/cuda symlink, observed in smg CI), the build dies with error: macro "__cudaLaunch" passed 2 arguments, but takes just 1. CUDA_HOME/include and the nvidia/cu*/include block already cover what nvcc needs, so drop the Triton block.

Verified inline-patched in smg-project/smg#1515 — once that CI is green the same change is safe to merge here.

@zhyncs
zhyncs requested a review from a team as a code owner May 25, 2026 09:17
``_resolve_include_dirs`` adds ``tokenspeed_triton/backends/nvidia/include``
and ``triton/backends/nvidia/include`` to nvcc's ``-I`` whenever the
matching ``cuda_runtime.h`` exists. Those Triton bundles ship a
CUDA-12-era ``crt/host_runtime.h`` whose ``__cudaLaunch`` is a 1-arg
macro; nvcc 13.0's generated ``cudafe1.stub.c`` calls it with 2 args.

When the host's system CUDA install is on nvcc's implicit search path
(e.g. a baked-in image) the system ``crt/host_runtime.h`` wins and
nothing breaks. On hosts where the system layout doesn't get
auto-picked (e.g. apt-installed CUDA 13 SDK exposed via a
``/usr/local/cuda`` symlink, observed downstream in smg CI), nvcc
falls through to the Triton bundles and the build fails with:

    error: macro "__cudaLaunch" passed 2 arguments, but takes just 1

The CUDA runtime headers nvcc needs are already covered by
``CUDA_HOME/include`` (from ``_cuda_toolkit_roots``) and the
``nvidia/cu*/include`` PyPI packages — folding the Triton bundles in
is gratuitous and unsafe, so drop the block entirely.

Signed-off-by: zhyncs <46627482+zhyncs@users.noreply.github.com>
@zhyncs
zhyncs force-pushed the zhyncs/kernel-drop-triton-include-path branch from 1e2d274 to ce9d989 Compare May 25, 2026 09:40
@lightseek-bot
lightseek-bot merged commit 5e145af into main May 25, 2026
2 checks passed
@lightseek-bot
lightseek-bot deleted the zhyncs/kernel-drop-triton-include-path branch May 25, 2026 09:41
zhyncs added a commit to chenht2022/smg that referenced this pull request May 25, 2026
- Drop the per-requirements-file installs and the wheel-cache scaffolding;
  let ``uv pip install -e ./python | tokenspeed-kernel/python | tokenspeed-scheduler/``
  resolve their own dependency metadata.
- Preseed ``setuptools wheel pybind11`` so the two ``--no-build-isolation``
  builds find their backend in the venv.
- Bump ``TOKENSPEED_REF`` to ``5e145afa`` (post lightseekorg/tokenspeed#247)
  so kernel build no longer fails on the stale Triton-bundled
  ``__cudaLaunch`` macro.

Signed-off-by: zhyncs <46627482+zhyncs@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants