launchers: point CUDA_HOME at the venv's CUDA 13 when the system nvcc is older - #185
Merged
Merged
Conversation
… is older vLLM 0.29's FlashInfer 0.6.18 JIT passes nvcc flags CUDA 12 does not know, so a native install with an older system nvcc dies on the first boot that compiles a FlashInfer kernel. FlashInfer takes CUDA_HOME before 'which nvcc' and pip already installed a CUDA 13 toolchain in the venv. Only when CUDA_HOME is unset and the nvcc on PATH is older than 13 or missing; the Docker image carries CUDA 13 and is untouched.
cpuchip
added a commit
to cpuchip/qwen38-27b-rtx3090
that referenced
this pull request
Sep 23, 2026
…vcc, syv-ai#171 tokenize contract test)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to the vLLM 0.29 pin (#148), for native installs.
The failure
0.29 ships FlashInfer 0.6.18, whose JIT passes nvcc flags CUDA 12 does not know. On a native install whose system
nvccis older than 13, the first boot that compiles a FlashInfer kernel dies:SPEC=mtp CTX=long(fp8 prefill) is the setup that hit it on the reference 3090, whose/usr/bin/nvccis 12.0. Upgraders from 0.28 meet it on their first boot after the pin, because 0.6.18 invalidates every cached FlashInfer kernel. The Docker image carries CUDA 13 and is unaffected.The fix
FlashInfer takes
CUDA_HOMEbeforewhich nvcc(flashinfer/jit/cpp_ext.py), and pip already installed a CUDA 13 toolchain in the venv (nvidia/cu13). Both launchers now setCUDA_HOMEto it, and say so at boot, only whenCUDA_HOMEis unset and thenvcconPATHis older than 13 or missing. An explicitCUDA_HOMEalways wins.Verification
Logic, over five cases (fake
nvccand venv): nvcc 12 + venv toolchain → set; nvcc 13 (the image) → untouched; no nvcc + venv toolchain → set; nvcc 12 + explicitCUDA_HOME→ kept; nvcc 12 + no venv toolchain → untouched.bash -nclean on both launchers.Live, on the reference 3090, main's 0.29 tree,
SPEC=mtp CTX=long, each arm with a fresh FlashInfer cache (FLASHINFER_WORKSPACE_BASEpointed at an empty dir) so the kernel really compiles:CUDA_HOME=/usr(system nvcc 12)nvcc fatal: Unknown option '--compress-mode=size', engine fails to start, 0 kernels compiledCUDA_HOMEunsetnvcc on PATH is 12 … CUDA_HOME=…/nvidia/cu13, 1 kernel compiled into the fresh cache, boots at the known 202,806-token pool,/health200