Skip to content

launchers: point CUDA_HOME at the venv's CUDA 13 when the system nvcc is older - #185

Merged
mhenrichsen merged 1 commit into
mainfrom
fix/nvcc-cu13-preflight
Sep 23, 2026
Merged

mhenrichsen merged 1 commit into
mainfrom
fix/nvcc-cu13-preflight

Conversation

@mhenrichsen

Copy link
Copy Markdown
Contributor

Follow-up to the vLLM 0.29 pin (#148), for native installs.

The failure

0.29 ships FlashInfer 0.6.18, whose JIT passes nvcc flags CUDA 12 does not know. On a native install whose system nvcc is older than 13, the first boot that compiles a FlashInfer kernel dies:

nvcc fatal   : Unknown option '--compress-mode=size'

SPEC=mtp CTX=long (fp8 prefill) is the setup that hit it on the reference 3090, whose /usr/bin/nvcc is 12.0. Upgraders from 0.28 meet it on their first boot after the pin, because 0.6.18 invalidates every cached FlashInfer kernel. The Docker image carries CUDA 13 and is unaffected.

The fix

FlashInfer takes CUDA_HOME before which nvcc (flashinfer/jit/cpp_ext.py), and pip already installed a CUDA 13 toolchain in the venv (nvidia/cu13). Both launchers now set CUDA_HOME to it, and say so at boot, only when CUDA_HOME is unset and the nvcc on PATH is older than 13 or missing. An explicit CUDA_HOME always wins.

Verification

Logic, over five cases (fake nvcc and venv): nvcc 12 + venv toolchain → set; nvcc 13 (the image) → untouched; no nvcc + venv toolchain → set; nvcc 12 + explicit CUDA_HOME → kept; nvcc 12 + no venv toolchain → untouched. bash -n clean on both launchers.

Live, on the reference 3090, main's 0.29 tree, SPEC=mtp CTX=long, each arm with a fresh FlashInfer cache (FLASHINFER_WORKSPACE_BASE pointed at an empty dir) so the kernel really compiles:

arm result
control: CUDA_HOME=/usr (system nvcc 12) nvcc fatal: Unknown option '--compress-mode=size', engine fails to start, 0 kernels compiled
this PR, CUDA_HOME unset launcher prints nvcc on PATH is 12 … CUDA_HOME=…/nvidia/cu13, 1 kernel compiled into the fresh cache, boots at the known 202,806-token pool, /health 200

… is older

vLLM 0.29's FlashInfer 0.6.18 JIT passes nvcc flags CUDA 12 does not know,
so a native install with an older system nvcc dies on the first boot that
compiles a FlashInfer kernel. FlashInfer takes CUDA_HOME before 'which nvcc'
and pip already installed a CUDA 13 toolchain in the venv. Only when
CUDA_HOME is unset and the nvcc on PATH is older than 13 or missing; the
Docker image carries CUDA 13 and is untouched.
@mhenrichsen
mhenrichsen merged commit fa97789 into main Sep 23, 2026
3 checks passed
@mhenrichsen
mhenrichsen deleted the fix/nvcc-cu13-preflight branch September 23, 2026 06:42
cpuchip added a commit to cpuchip/qwen38-27b-rtx3090 that referenced this pull request Sep 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant