Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
77 commits
Select commit Hold shift + click to select a range
8eafc8e
Upgrade dependencies for dlfw 26.03 stack
EmmaQiaoCh Apr 1, 2026
e076851
Update image tags
EmmaQiaoCh Apr 1, 2026
9623ed4
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Apr 1, 2026
8162d6d
FFix a clang missing-braces error
EmmaQiaoCh Apr 2, 2026
279844e
Fix again
EmmaQiaoCh Apr 2, 2026
02c631a
Fix build error again
EmmaQiaoCh Apr 3, 2026
eba68f4
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Apr 3, 2026
6390b18
Upgrade onnxscript to support torch 2.11
EmmaQiaoCh Apr 7, 2026
2032170
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Apr 7, 2026
8eacfce
Testing for the reservation
EmmaQiaoCh Apr 13, 2026
73cb947
Add workaround due to API incompability between public torch 2.11 and…
EmmaQiaoCh Apr 13, 2026
e265cd9
Test for the H100 node which has new cuda driver
EmmaQiaoCh Apr 15, 2026
9c25524
Correct a typo
EmmaQiaoCh Apr 15, 2026
ba7ea4e
Fix the h100 selector condition
EmmaQiaoCh Apr 15, 2026
25cf6d5
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Apr 16, 2026
fc49d6a
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Apr 17, 2026
31ea317
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Apr 19, 2026
586686d
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Apr 21, 2026
5eeb6a8
Add back the notes which may still need
EmmaQiaoCh Apr 24, 2026
c972181
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Apr 24, 2026
394328b
Fix some typo
EmmaQiaoCh Apr 27, 2026
e037d3a
Test for specific node which has upgraded cuda driver
EmmaQiaoCh Apr 27, 2026
b2d0e1c
Specified the test node
EmmaQiaoCh Apr 28, 2026
799a420
Fix affinity
EmmaQiaoCh Apr 28, 2026
af853e1
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Apr 29, 2026
17552a4
Update for 26.04 base image
EmmaQiaoCh Apr 29, 2026
b346db2
Update image tags and remove testing code for dgx_h100
EmmaQiaoCh Apr 29, 2026
332abea
Update multi-node mpi
EmmaQiaoCh Apr 30, 2026
3cfc932
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Apr 30, 2026
8766b9d
Remove the work around for installation of constraints
EmmaQiaoCh May 6, 2026
6be67de
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh May 6, 2026
8e674da
Use pip cache
EmmaQiaoCh May 6, 2026
7a79475
Fix nvshmem header to use them in current dir
EmmaQiaoCh May 7, 2026
ac5ef1c
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh May 7, 2026
d1429e5
Fix error for '-fclang-abi-compat=17'
EmmaQiaoCh May 8, 2026
d28ebb5
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh May 8, 2026
899b623
Add 2 dependencies which are removed from pytorch image,fix another c…
EmmaQiaoCh May 9, 2026
34ca9d3
Waive failed cases
EmmaQiaoCh May 11, 2026
4388bdf
Remove waives for perf cases after check with dev
EmmaQiaoCh May 11, 2026
e8dd87e
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh May 11, 2026
9b9121e
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh May 12, 2026
30f608f
Remove duplicated waives
EmmaQiaoCh May 12, 2026
d1a038e
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh May 14, 2026
ec56066
Rebuild image due to changes from main
EmmaQiaoCh May 14, 2026
51d3e9e
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh May 14, 2026
a9e44b2
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh May 18, 2026
5fc0f9d
Apply dev's fix for nvbug 6162517
EmmaQiaoCh May 20, 2026
13b7e24
Waive some failed cases in new base
EmmaQiaoCh May 21, 2026
a3a43e0
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh May 21, 2026
e13cc4e
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh May 24, 2026
a4c37c1
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh May 26, 2026
0e51667
Fix for verl tests from dev
EmmaQiaoCh May 26, 2026
3c9c552
Remove 2 waives
EmmaQiaoCh May 27, 2026
9e4e14d
Upgrade onnx in requirements
EmmaQiaoCh May 27, 2026
e660300
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh May 28, 2026
77e5421
Waive failed stage and cases
EmmaQiaoCh May 28, 2026
dd5a26b
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Jun 1, 2026
361a021
Merge branch 'NVIDIA:main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Jun 1, 2026
fa43d84
Apply the fix from nv-auto-deploy@6370238
EmmaQiaoCh Jun 2, 2026
64c66d9
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Jun 2, 2026
270cf96
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Jun 3, 2026
a105a8b
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Jun 4, 2026
e92a2e2
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Jun 4, 2026
5432e7a
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Jun 5, 2026
01c11cf
Update new image tags and unwaive a case
EmmaQiaoCh Jun 5, 2026
bc63aa2
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Jun 6, 2026
1a8eac2
Change back mpi for perf multi-node
EmmaQiaoCh Jun 8, 2026
2df7e5b
Update Docker image in x86 sanity check config
EmmaQiaoCh Jun 8, 2026
c23401a
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Jun 8, 2026
37af6de
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Jun 9, 2026
efff321
Update image tags
EmmaQiaoCh Jun 9, 2026
5e8ba12
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Jun 9, 2026
23fc5e6
Merge branch 'main' into emma/upgrade_dlfw_2603
EmmaQiaoCh Jun 10, 2026
d371e05
[TRTLLM-11715][infra] Fix MPI/PMI plumbing for DLFW 26.03 OpenMPI build
chenfeiz0326 Jun 11, 2026
7c62be5
[None][infra] Reset waives.txt to origin/main
chenfeiz0326 Jun 11, 2026
60d0771
Merge branch 'main' into chenfeiz/fix-emma-mpi-pmix-26-03
chenfeiz0326 Jun 11, 2026
6fe22ab
[TRTLLM-11715][infra] Force PMIX_MCA_gds=hash in multi-node slurm_run
chenfeiz0326 Jun 15, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,8 +8,8 @@ TensorRT LLM
[![Ask DeepWiki](https://deepwiki.com/badge.svg)](https://deepwiki.com/NVIDIA/TensorRT-LLM)
[![python](https://img.shields.io/badge/python-3.12-green)](https://www.python.org/downloads/release/python-3123/)
[![python](https://img.shields.io/badge/python-3.10-green)](https://www.python.org/downloads/release/python-31012/)
[![cuda](https://img.shields.io/badge/cuda-13.1.1-green)](https://developer.nvidia.com/cuda-downloads)
[![torch](https://img.shields.io/badge/torch-2.10.0-green)](https://pytorch.org)
[![cuda](https://img.shields.io/badge/cuda-13.2.1-green)](https://developer.nvidia.com/cuda-downloads)
[![torch](https://img.shields.io/badge/torch-2.11.0-green)](https://pytorch.org)
[![version](https://img.shields.io/badge/release-1.3.0rc18-green)](https://github.com/NVIDIA/TensorRT-LLM/blob/main/tensorrt_llm/version.py)
[![license](https://img.shields.io/badge/license-Apache%202-blue)](https://github.com/NVIDIA/TensorRT-LLM/blob/main/LICENSE)

Expand Down
6 changes: 0 additions & 6 deletions constraints.txt
Original file line number Diff line number Diff line change
@@ -1,11 +1,5 @@
# These vulnerabilities were inherited from the base image (pytorch:25.12-py3) and should be removed when the base image
# is updated.
# WAR against https://github.com/advisories/GHSA-8rrh-rw8j-w5fx
wheel>=0.46.2
# WAR against https://github.com/advisories/GHSA-qjxf-f2mg-c6mc
tornado>=6.5.5
# WAR against https://github.com/advisories/GHSA-3936-cmfr-pm3m
black>=26.3.1
# Upgrade base image nvidia-cutlass-dsl 4.3.5 to 4.4.2
nvidia-cutlass-dsl>=4.4.2
# The `nvidia-cutlass-dsl` package does not pin numpy at all, which can be problematic in certain CI
Expand Down
5 changes: 1 addition & 4 deletions cpp/include/tensorrt_llm/runtime/virtualMemory.h
Original file line number Diff line number Diff line change
Expand Up @@ -505,10 +505,7 @@ class CudaVirtualMemoryAllocator
{
std::size_t gpuAlignment = 1;
CUmemAllocationProp const prop{CU_MEM_ALLOCATION_TYPE_PINNED, CU_MEM_HANDLE_TYPE_NONE,
{
CU_MEM_LOCATION_TYPE_DEVICE,
device,
}};
CUmemLocation{CU_MEM_LOCATION_TYPE_DEVICE, {device}}};
TLLM_CU_CHECK(
cuMemGetAllocationGranularity(&gpuAlignment, &prop, CU_MEM_ALLOC_GRANULARITY_RECOMMENDED));
alignment = std::lcm(getpagesize(), gpuAlignment);
Expand Down
15 changes: 15 additions & 0 deletions cpp/tensorrt_llm/deep_ep/CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -120,6 +120,15 @@ if(NOT CMAKE_CXX_COMPILER_ID STREQUAL "GNU")
set(CMAKE_C_COMPILER gcc)
set(CMAKE_CXX_COMPILER g++)
set(CMAKE_CUDA_HOST_COMPILER g++)
# PyTorch's cmake/public/cuda.cmake (loaded transitively by
# find_package(Torch)) appends -Xcompiler=-fclang-abi-compat=17 to
# CMAKE_CUDA_FLAGS whenever the parent build is configured with Clang>=18 (see
# pytorch PR #175233). Since this subdirectory falls back to GCC for NVSHMEM
# compatibility, that Clang-only flag would be forwarded to g++ via `nvcc
# -ccbin=g++` and abort the build with: g++: error: unrecognized command-line
# option '-fclang-abi-compat=17'
string(REGEX REPLACE "-Xcompiler=-fclang-abi-compat=[0-9]+" ""
CMAKE_CUDA_FLAGS "${CMAKE_CUDA_FLAGS}")
endif()

# Add nvshmem external project
Expand Down Expand Up @@ -204,6 +213,12 @@ target_compile_options(
target_compile_definitions(
deep_ep_cpp_tllm PRIVATE DISABLE_AGGRESSIVE_PTX_INSTRS
TORCH_EXTENSION_NAME=deep_ep_cpp_tllm)
# Newer CUDA containers provide NVSHMEM headers in the default CUDA include
# directory. DeepEP must compile against the vendored NVSHMEM headers because it
# links the vendored NVSHMEM static library below.
target_include_directories(
deep_ep_cpp_tllm BEFORE
PRIVATE ${CMAKE_CURRENT_BINARY_DIR}/nvshmem-build/src/include)
target_link_libraries(
deep_ep_cpp_tllm PRIVATE nvshmem_project::nvshmem ${TORCH_LIBRARIES}
${TORCH_PYTHON_LIB})
Expand Down
9 changes: 9 additions & 0 deletions cpp/tensorrt_llm/flash_mla/CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -44,6 +44,15 @@ if(CMAKE_SYSTEM_PROCESSOR MATCHES "aarch64|arm64" AND CMAKE_CXX_COMPILER_ID
set(CMAKE_CUDA_HOST_COMPILER ${GCC_EXECUTABLE})
message(
STATUS "FlashMLA: Using GCC at ${GCC_EXECUTABLE} for CUDA compilation")
# PyTorch's cmake/public/cuda.cmake (loaded transitively by
# find_package(Torch)) appends -Xcompiler=-fclang-abi-compat=17 to
# CMAKE_CUDA_FLAGS whenever the parent build is configured with Clang>=18 (see
# pytorch PR #175233). Since CUDA host compilation here falls back to GCC,
# that Clang-only flag would be forwarded to g++ via `nvcc -ccbin=g++` and
# abort the build with: g++: error: unrecognized command-line option
# '-fclang-abi-compat=17'
string(REGEX REPLACE "-Xcompiler=-fclang-abi-compat=[0-9]+" ""
CMAKE_CUDA_FLAGS "${CMAKE_CUDA_FLAGS}")
endif()

# Check CUDA version and architecture support
Expand Down
11 changes: 2 additions & 9 deletions cpp/tensorrt_llm/runtime/virtualMemory.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -344,11 +344,7 @@ void CudaVirtualMemoryAllocator::allocate(Pointer* ptr, std::size_t n, int devic

CUDAVirtualMemoryChunk::Configurators configurators;
configurators.push_back(std::make_unique<UnicastConfigurator>(address, alignedSize,
CUmemAccessDesc{{
CU_MEM_LOCATION_TYPE_DEVICE,
device,
},
CU_MEM_ACCESS_FLAGS_PROT_READWRITE}));
CUmemAccessDesc{CUmemLocation{CU_MEM_LOCATION_TYPE_DEVICE, {device}}, CU_MEM_ACCESS_FLAGS_PROT_READWRITE}));

switch (mConfig->mMode)
{
Expand All @@ -368,10 +364,7 @@ void CudaVirtualMemoryAllocator::allocate(Pointer* ptr, std::size_t n, int devic

mConfig->mManager.add(address, mConfig->mTag,
std::make_unique<LocalCreator<>>(CUmemAllocationProp{CU_MEM_ALLOCATION_TYPE_PINNED, CU_MEM_HANDLE_TYPE_NONE,
{
CU_MEM_LOCATION_TYPE_DEVICE,
device,
}},
CUmemLocation{CU_MEM_LOCATION_TYPE_DEVICE, {device}}},
alignedSize),
std::move(configurators));

Expand Down
4 changes: 2 additions & 2 deletions cpp/tests/unit_tests/common/cudaDriverWrapperTest.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,7 @@ TEST(TestCudaDriverWrapper, TllmCuCheckFailingWithValidParametersDoesNotThrow)
CUmemAllocationHandleType::CU_MEM_HANDLE_TYPE_NONE,
CUmemLocation{
CUmemLocationType::CU_MEM_LOCATION_TYPE_DEVICE,
0,
{0},
},
nullptr};
auto const granularity = tensorrt_llm::common::getAllocationGranularity();
Expand All @@ -51,7 +51,7 @@ TEST(TestCudaDriverWrapper, TllmCuCheckFailingWithInvalidParametersThrows)
CUmemAllocationHandleType::CU_MEM_HANDLE_TYPE_NONE,
CUmemLocation{
CUmemLocationType::CU_MEM_LOCATION_TYPE_DEVICE,
0,
{0},
},
nullptr};
ASSERT_THROW(TLLM_CU_CHECK(cuMemCreate(&handle, -1, &prop, 0ULL)), tensorrt_llm::common::TllmException);
Expand Down
45 changes: 12 additions & 33 deletions cpp/tests/unit_tests/runtime/virtualMemoryTest.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -142,19 +142,12 @@ TEST_F(VirtualMemoryTest, TestBasic)

CUDAVirtualMemoryChunk::CreatorPtr creator
= std::make_unique<LocalCreator<>>(CUmemAllocationProp{CU_MEM_ALLOCATION_TYPE_PINNED, CU_MEM_HANDLE_TYPE_NONE,
{
CU_MEM_LOCATION_TYPE_DEVICE,
0,
}},
CUmemLocation{CU_MEM_LOCATION_TYPE_DEVICE, {0}}},
size);

CUDAVirtualMemoryChunk::Configurators configurators;
configurators.push_back(std::make_unique<UnicastConfigurator>(address, size,
CUmemAccessDesc{{
CU_MEM_LOCATION_TYPE_DEVICE,
0,
},
CU_MEM_ACCESS_FLAGS_PROT_READWRITE}));
CUmemAccessDesc{CUmemLocation{CU_MEM_LOCATION_TYPE_DEVICE, {0}}, CU_MEM_ACCESS_FLAGS_PROT_READWRITE}));

CUDAVirtualMemoryChunk vm(std::move(creator), std::move(configurators));
ASSERT_EQ(vm.status(), CUDAVirtualMemoryChunk::RELEASED);
Expand Down Expand Up @@ -212,19 +205,12 @@ TEST_P(VirtualMemoryOffloadConfigurator, Test)

CUDAVirtualMemoryChunk::CreatorPtr creator
= std::make_unique<LocalCreator<>>(CUmemAllocationProp{CU_MEM_ALLOCATION_TYPE_PINNED, CU_MEM_HANDLE_TYPE_NONE,
{
CU_MEM_LOCATION_TYPE_DEVICE,
0,
}},
CUmemLocation{CU_MEM_LOCATION_TYPE_DEVICE, {0}}},
size);

CUDAVirtualMemoryChunk::Configurators configurators;
configurators.push_back(std::make_unique<UnicastConfigurator>(address, size,
CUmemAccessDesc{{
CU_MEM_LOCATION_TYPE_DEVICE,
0,
},
CU_MEM_ACCESS_FLAGS_PROT_READWRITE}));
CUmemAccessDesc{CUmemLocation{CU_MEM_LOCATION_TYPE_DEVICE, {0}}, CU_MEM_ACCESS_FLAGS_PROT_READWRITE}));
configurators.push_back(std::make_unique<OffloadConfigurator>(address, size, backType, stream.get(), false));

CUDAVirtualMemoryChunk vm(std::move(creator), std::move(configurators));
Expand Down Expand Up @@ -611,14 +597,14 @@ TEST_F(VirtualMemoryTest, TestFacilities)
{

// Create original CUDAVirtualMemoryChunk
CUDAVirtualMemoryChunk::CreatorPtr creator
= std::make_unique<LocalCreator<>>(CUmemAllocationProp{CU_MEM_ALLOCATION_TYPE_PINNED,
CU_MEM_HANDLE_TYPE_NONE, {CU_MEM_LOCATION_TYPE_DEVICE, 0}},
size);
CUDAVirtualMemoryChunk::CreatorPtr creator = std::make_unique<LocalCreator<>>(
CUmemAllocationProp{CU_MEM_ALLOCATION_TYPE_PINNED, CU_MEM_HANDLE_TYPE_NONE,
CUmemLocation{CU_MEM_LOCATION_TYPE_DEVICE, {0}}},
size);

CUDAVirtualMemoryChunk::Configurators configurators;
configurators.push_back(std::make_unique<UnicastConfigurator>(
address, size, CUmemAccessDesc{{CU_MEM_LOCATION_TYPE_DEVICE, 0}, CU_MEM_ACCESS_FLAGS_PROT_READWRITE}));
configurators.push_back(std::make_unique<UnicastConfigurator>(address, size,
CUmemAccessDesc{CUmemLocation{CU_MEM_LOCATION_TYPE_DEVICE, {0}}, CU_MEM_ACCESS_FLAGS_PROT_READWRITE}));

CUDAVirtualMemoryChunk original(std::move(creator), std::move(configurators));
original.materialize();
Expand Down Expand Up @@ -977,19 +963,12 @@ TEST_F(VirtualMemoryManagerTest, TestBasic)

CUDAVirtualMemoryChunk::CreatorPtr creator
= std::make_unique<LocalCreator<>>(CUmemAllocationProp{CU_MEM_ALLOCATION_TYPE_PINNED, CU_MEM_HANDLE_TYPE_NONE,
{
CU_MEM_LOCATION_TYPE_DEVICE,
0,
}},
CUmemLocation{CU_MEM_LOCATION_TYPE_DEVICE, {0}}},
size);

CUDAVirtualMemoryChunk::Configurators configurators;
configurators.push_back(std::make_unique<UnicastConfigurator>(address, size,
CUmemAccessDesc{{
CU_MEM_LOCATION_TYPE_DEVICE,
0,
},
CU_MEM_ACCESS_FLAGS_PROT_READWRITE}));
CUmemAccessDesc{CUmemLocation{CU_MEM_LOCATION_TYPE_DEVICE, {0}}, CU_MEM_ACCESS_FLAGS_PROT_READWRITE}));

auto memoryBegin = getCurrentProcessMemoryInfo();

Expand Down
4 changes: 2 additions & 2 deletions docker/Dockerfile.multi
Original file line number Diff line number Diff line change
@@ -1,8 +1,8 @@
# Multi-stage Dockerfile
ARG BASE_IMAGE=nvcr.io/nvidia/pytorch
ARG TRITON_IMAGE=nvcr.io/nvidia/tritonserver
ARG BASE_TAG=26.02-py3
ARG TRITON_BASE_TAG=26.02-py3
ARG BASE_TAG=26.04-py3
ARG TRITON_BASE_TAG=26.04-py3
ARG DEVEL_IMAGE=devel

FROM ${BASE_IMAGE}:${BASE_TAG} AS base
Expand Down
8 changes: 4 additions & 4 deletions docker/Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -202,21 +202,21 @@ jenkins-rockylinux8_%: PYTHON_VERSION_TAG_ID = $(if $(findstring 3.12,${PYTHON_V
jenkins-rockylinux8_%: IMAGE_WITH_TAG = $(shell . ../jenkins/current_image_tags.properties && echo $$LLM_ROCKYLINUX8_${PYTHON_VERSION_TAG_ID}_DOCKER_IMAGE)
jenkins-rockylinux8_%: STAGE = tritondevel
jenkins-rockylinux8_%: BASE_IMAGE = nvcr.io/nvidia/cuda
jenkins-rockylinux8_%: BASE_TAG = 13.1.1-devel-rockylinux8
jenkins-rockylinux8_%: BASE_TAG = 13.2.1-devel-rockylinux8

rockylinux8_%: STAGE = tritondevel
rockylinux8_%: BASE_IMAGE = nvcr.io/nvidia/cuda
rockylinux8_%: BASE_TAG = 13.1.1-devel-rockylinux8
rockylinux8_%: BASE_TAG = 13.2.1-devel-rockylinux8

# For x86_64
ubuntu22_%: STAGE = tritondevel
ubuntu22_%: BASE_IMAGE = nvcr.io/nvidia/cuda
ubuntu22_%: BASE_TAG = 13.1.1-devel-ubuntu22.04
ubuntu22_%: BASE_TAG = 13.2.1-devel-ubuntu22.04

# For x86_64 and aarch64
ubuntu24_%: STAGE = tritondevel
ubuntu24_%: BASE_IMAGE = nvcr.io/nvidia/cuda
ubuntu24_%: BASE_TAG = 13.1.1-devel-ubuntu24.04
ubuntu24_%: BASE_TAG = 13.2.1-devel-ubuntu24.04

trtllm_%: STAGE = release
trtllm_%: PUSH_TO_STAGING := 0
Expand Down
2 changes: 1 addition & 1 deletion docker/common/install_cuda_toolkit.sh
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ set -ex
# This script is used for reinstalling CUDA on Rocky Linux 8 with the run file.
# CUDA version is usually aligned with the latest NGC CUDA image tag.
# Only use when public CUDA image is not ready.
CUDA_VER="13.1.1_590.48.01"
CUDA_VER="13.2.1_595.58.03"
CUDA_VER_SHORT="${CUDA_VER%_*}"

NVCC_VERSION_OUTPUT=$(nvcc --version)
Expand Down
4 changes: 2 additions & 2 deletions docker/common/install_pytorch.sh
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,8 @@ set -ex

# Use latest stable version from https://pypi.org/project/torch/#history
# and closest to the version specified in
# https://docs.nvidia.com/deeplearning/frameworks/pytorch-release-notes/rel-26-02.html#rel-26-02
TORCH_VERSION="2.10.0"
# https://docs.nvidia.com/deeplearning/frameworks/pytorch-release-notes/rel-26-04.html#rel-26-04
TORCH_VERSION="2.11.0"
SYSTEM_ID=$(grep -oP '(?<=^ID=).+' /etc/os-release | tr -d '"')

prepare_environment() {
Expand Down
18 changes: 9 additions & 9 deletions docker/common/install_tensorrt.sh
Original file line number Diff line number Diff line change
Expand Up @@ -2,20 +2,20 @@

set -ex

TRT_VER="10.15.1.29"
TRT_VER="10.16.1.11"
# Align with the pre-installed cuDNN / cuBLAS / NCCL versions from
# https://docs.nvidia.com/deeplearning/frameworks/pytorch-release-notes/rel-26-02.html#rel-26-02
CUDA_VER="13.1" # 13.1.1
# https://docs.nvidia.com/deeplearning/frameworks/pytorch-release-notes/rel-26-04.html#rel-26-04
CUDA_VER="13.2" # 13.2.1
# Keep the installation for cuDNN if users want to install PyTorch with source codes.
# PyTorch 2.x can compile with cuDNN v9.
CUDNN_VER="9.19.0.56-1"
NCCL_VER="2.29.2-1+cuda13.1"
CUBLAS_VER="13.2.1.1-1"
CUDNN_VER="9.21.0.82-1"
NCCL_VER="2.29.7-1+cuda13.2"
CUBLAS_VER="13.4.0.1-1"
# Align with the pre-installed CUDA / NVCC / NVRTC versions from
# https://docs.nvidia.com/cuda/cuda-toolkit-release-notes/index.html
NVRTC_VER="13.1.115-1"
CUDA_RUNTIME="13.1.80-1"
CUDA_DRIVER_VERSION="590.48.01-1.el8"
NVRTC_VER="13.2.78-1"
CUDA_RUNTIME="13.2.75-1"
CUDA_DRIVER_VERSION="595.58.03-1.el8"

for i in "$@"; do
case $i in
Expand Down
4 changes: 2 additions & 2 deletions docs/source/legacy/reference/support-matrix.md
Original file line number Diff line number Diff line change
Expand Up @@ -158,9 +158,9 @@ The following table shows the supported software for TensorRT-LLM.
* -
- Software Compatibility
* - Container
- [26.02](https://docs.nvidia.com/deeplearning/frameworks/support-matrix/index.html)
- [26.04](https://docs.nvidia.com/deeplearning/frameworks/support-matrix/index.html)
* - TensorRT
- [10.14](https://docs.nvidia.com/deeplearning/tensorrt/release-notes/index.html)
- [10.16](https://docs.nvidia.com/deeplearning/tensorrt/release-notes/index.html)
* - Precision
-
- Blackwell (SM100/SM103/SM120) - FP32, FP16, BF16, FP8, FP4, INT8, INT4
Expand Down
2 changes: 1 addition & 1 deletion jenkins/Build.groovy
Original file line number Diff line number Diff line change
Expand Up @@ -408,7 +408,7 @@ def runLLMBuild(pipeline, buildFlags, tarName, is_linux_x86_64)
def llmPath = sh (script: "realpath ${LLM_ROOT}",returnStdout: true).trim()
// TODO: Remove after the cmake version is upgraded to 3.31.8
// Get triton tag from docker/dockerfile.multi
def tritonShortTag = "r26.02"
def tritonShortTag = "r26.04"
sh "cd ${LLM_ROOT}/triton_backend/inflight_batcher_llm && mkdir build && cd build && cmake .. -DTRTLLM_DIR=${llmPath} -DTRITON_COMMON_REPO_TAG=${tritonShortTag} -DTRITON_CORE_REPO_TAG=${tritonShortTag} -DTRITON_THIRD_PARTY_REPO_TAG=${tritonShortTag} -DTRITON_BACKEND_REPO_TAG=${tritonShortTag} -DUSE_CXX11_ABI=ON && make -j${buildJobs} install"

// Step 3: packaging wheels into tarfile
Expand Down
Loading
Loading