diff --git a/examples/inference/README.md b/examples/inference/README.md index fe501635f..de62b318a 100644 --- a/examples/inference/README.md +++ b/examples/inference/README.md @@ -13,7 +13,8 @@ Framework-centric inference engine examples, organized by serving engine. | [`sglang`](./sglang) | [`dsv4pro-b300-single-node`](./sglang/dsv4pro-b300-single-node) | DeepSeek V4 Pro unified (non-PD) serving on a single B300 node | | [`sglang`](./sglang) | [`dsv4flash-b300-intra-3p1d`](./sglang/dsv4flash-b300-intra-3p1d) | DeepSeek-V4-Flash with intra-node 3-prefill / 1-decode disaggregation (tp=2 each) on a single B300 node | | [`sglang`](./sglang) | [`glm5.2-b300-tp2-dp4`](./sglang/glm5.2-b300-tp2-dp4) | GLM-5.2 (NVFP4) as 4× independent tp=2 engines behind an SGLang router on a single B300 node | +| [`tensorrt-llm`](./tensorrt-llm) | [`nccl-ep-efa`](./tensorrt-llm/nccl-ep-efa) | `Qwen/Qwen3-30B-A3B` MoE with TensorRT-LLM's `NcclEP` wide-EP dispatch/combine (ep_size=8) over EFA via the NCCL-GIN CPU-proxy on 2× `p5en.48xlarge` (H200) — EKS | -More engines (TRT-LLM, NIM, Ray Serve, …) are planned, including +More engines (NIM, Ray Serve, …) are planned, including content to be merged from [`aws-samples/awsome-inference`](https://github.com/aws-samples/awsome-inference) (see issue #1056). diff --git a/examples/inference/tensorrt-llm/README.md b/examples/inference/tensorrt-llm/README.md new file mode 100644 index 000000000..c28e370b6 --- /dev/null +++ b/examples/inference/tensorrt-llm/README.md @@ -0,0 +1,22 @@ + + +# TensorRT-LLM test cases + +[TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM) is NVIDIA's open-source +inference engine for large language models on NVIDIA GPUs. The samples in this +directory deploy TensorRT-LLM on AWS with high-performance EFA networking and +expert-parallel MoE all-to-all. + +## Available test cases + +| Test case | Orchestrator | Description | +| --- | --- | --- | +| [`nccl-ep-efa`](./nccl-ep-efa) | Kubernetes (2-node) | Wide-EP MoE dispatch/combine via TRT-LLM's **`NcclEP`** backend (`nccl.ep` / `libnccl_ep` — NOT the `deep_ep` package) over **AWS EFA**, using aws-ofi-nccl with **GIN** (GPU-Initiated Networking) CPU-proxy. Image built NGC-from-scratch from public sources; the recipe runs image build → transport smoke test → served `/v1/chat/completions` → concurrency benchmark. Validated on `p5en.48xlarge` (H200). | + +For kernel-level expert-parallelism dispatch/combine benchmarks over EFA — +including a DeepEP V2 benchmark on the same NCCL-GIN substrate this test case +uses — see +[`micro-benchmarks/expert-parallelism`](../../../micro-benchmarks/expert-parallelism). diff --git a/examples/inference/tensorrt-llm/nccl-ep-efa/.gitignore b/examples/inference/tensorrt-llm/nccl-ep-efa/.gitignore new file mode 100644 index 000000000..d8f533dc2 --- /dev/null +++ b/examples/inference/tensorrt-llm/nccl-ep-efa/.gitignore @@ -0,0 +1,4 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. SPDX-License-Identifier: MIT-0 +setup/env_vars +benchmarks/raw/ +*.log diff --git a/examples/inference/tensorrt-llm/nccl-ep-efa/Dockerfile b/examples/inference/tensorrt-llm/nccl-ep-efa/Dockerfile new file mode 100644 index 000000000..2972da951 --- /dev/null +++ b/examples/inference/tensorrt-llm/nccl-ep-efa/Dockerfile @@ -0,0 +1,186 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. SPDX-License-Identifier: MIT-0 +# +# TensorRT-LLM + NcclEP MoE all-to-all over AWS EFA (NCCL-GIN CPU-proxy). +# TRT-LLM's NcclEP backend rides `nccl.ep` (nccl4py), which needs NCCL >= 2.30.4 and a +# GIN-capable network plugin — neither of which the NGC TRT-LLM container ships. This image +# adds, from PUBLIC sources only: EFA userspace, gdrcopy, the pinned NCCL + nccl4py pair, +# and aws-ofi-nccl's GIN plugin (built by setup_trtllm_nccl_ep_efa.sh — COPY'd, not curled, +# so it is in-tree + reviewable). +# +# setup/build-push.sh builds + pushes ${REGISTRY}/${IMAGE_NAME}:${IMAGE_TAG} (setup/env_vars); +# manual equivalent: DOCKER_BUILDKIT=1 docker build -t /trtllm-nccl-ep-efa: . +# +# Base pin rationale (GA-over-prerelease exception, documented): GA v1.2.1 does NOT contain +# the NcclEP backend at all (tensorrt_llm/_torch/modules/fused_moe/nccl_ep_utils.py is absent +# at that tag), so a 1.3.0rc pin is REQUIRED, not a preference. 1.3.0rc24 is the first NGC +# release container whose factory ships the NCCL_EP arm natively (TRTLLM_FORCE_COMM_METHOD); +# re-test and re-pin when a GA that carries the backend appears. +ARG TRTLLM_BASE=nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc24 +FROM ${TRTLLM_BASE} +ARG TRTLLM_BASE # re-declare: pre-FROM ARGs go out of scope after FROM (used in Layer 6 diagnostics) + +LABEL org.opencontainers.image.description="TensorRT-LLM + NcclEP MoE all-to-all over AWS EFA (NCCL-GIN CPU-proxy)" +LABEL org.opencontainers.image.licenses="MIT-0" +LABEL org.opencontainers.image.source="https://github.com/awslabs/awsome-distributed-ai" +ENV DEBIAN_FRONTEND=noninteractive +SHELL ["/bin/bash", "-c"] +USER root + +# ---- Layer 1: system + build deps ------------------------------------------- +# libevent-{core,pthreads} are prrte-aws deps the EFA installer needs; do NOT purge +# /var/lib/apt/lists here — the installer below runs its own apt-get install and fails +# with "held broken packages" against empty lists. Purged at the end of Layer 2 instead. +RUN apt-get update && apt-get install -y --no-install-recommends \ + build-essential autoconf automake libtool pkg-config git curl wget ca-certificates \ + libnuma-dev libhwloc-dev libudev-dev \ + libevent-core-2.1-7t64 libevent-pthreads-2.1-7t64 \ + pciutils environment-modules tcl udev dmidecode ethtool iproute2 kmod + +# ---- Layer 2: AWS EFA (public installer) ---- +# 1.50.0 is the current EFA userspace and what the sibling deepep-v2-benchmark builds; its +# tarball resolves (verified) and it ships aws-ofi-nccl 1.21.1 in-box, lining this sample up +# with the rest of the repo. BUILD-PIN only: the 2026-08-07 correctness E2E (real +# trtllm-serve HTTP-200-correct + 16-rank cross-node dispatch/combine, IMA=0, efa-direct on +# every rank) ran on 1.48.0 — this assembly is not yet cluster-re-measured on 1.50.0, which +# is exactly what recipe/verify-image.sh + run-kernel-test.sh exist to re-verify on-cluster. +# --disable-ngc: the NGC base trips the installer's NGC auto-detect, which would silently +# reroute to the libnccl-ofi-ngc path — we build aws-ofi-nccl from source ourselves +# (Layer 5), same explicit choice as the sibling vllm/dsv3-uccl-nixl sample (which also +# passes --disable-ngc alone). Do NOT add --disable-build-ngc: that flag existed in +# aws-efa-installer 1.48.0 but was REMOVED in 1.49.0+ (the installer no longer auto-detects +# the build-ngc use case, so the flag became unnecessary and getopt now rejects it — an +# unknown long-opt makes efa_installer.sh print usage and exit 1, failing this layer). The +# 2026-08-07 E2E ran on 1.48.0 where both flags were valid; the 1.50.0 bump made the second +# one a build error, which is why it is dropped here. +ARG EFA_INSTALLER_VER=1.50.0 +RUN apt-get update \ + && curl -fsSL https://efa-installer.amazonaws.com/aws-efa-installer-${EFA_INSTALLER_VER}.tar.gz | tar -xzf - -C /tmp \ + && cd /tmp/aws-efa-installer \ + && ./efa_installer.sh -y --skip-kmod --skip-limit-conf --no-verify --disable-ngc \ + && echo "${EFA_INSTALLER_VER}" > /opt/efa-installer.version \ + && rm -rf /tmp/aws-efa-installer /var/lib/apt/lists/* +# Add EFA userspace (libfabric + efa provider) to PATH/LD, but deliberately NOT the EFA +# installer's bundled OpenMPI (/opt/amazon/openmpi). The NGC TRT-LLM base ships HPC-X OpenMPI +# (ldconfig default /opt/hpcx/ompi) and tensorrt_llm's MPI_Init resolves libopen-pal through it. +# The EFA installer's OpenMPI 4.1.7 ships a libopen-pal.so.40 that does NOT export +# opal_libevent2022_event_assign; prepending /opt/amazon/openmpi/lib here shadows the HPC-X +# libopen-pal under HPC-X's own mca_ess_hnp.so plugin → undefined symbol → MPI_Init aborts on +# `import tensorrt_llm`. Neither recipe uses EFA's mpirun (serve.sh is single-node; the probe +# uses torchrun), so the EFA OpenMPI is unnecessary and its lib on LD is actively harmful. +ENV PATH=/opt/amazon/efa/bin:$PATH +ENV LD_LIBRARY_PATH=/opt/amazon/efa/lib:${LD_LIBRARY_PATH:-} + +# ---- Layer 3: gdrcopy userspace (PUBLIC: github.com/NVIDIA/gdrcopy) ---- +# aws-ofi-nccl's GIN path REQUIRES gdrapi.h at configure time — without it the plugin +# compiles with "GDRCopy support not available", nccl_ofi_gin_init fails at serve time, +# and NcclEP's group creation dies. gdrcopy v2.5.2 == commit c91ad9f: pin the commit, not +# the tag (a bare tag is a moving ref upstream can re-point). +ARG GDRCOPY_SHA=c91ad9f178e5fb729fc5b6dc62a77c3bb364d6c9 +RUN git clone https://github.com/NVIDIA/gdrcopy.git /tmp/gdrcopy \ + && cd /tmp/gdrcopy && git fetch origin ${GDRCOPY_SHA} && git checkout ${GDRCOPY_SHA} \ + && make prefix=/usr/local lib lib_install && ldconfig \ + && rm -rf /tmp/gdrcopy + +# ---- Layer 4: NCCL 2.30.4 (over the container's baked NCCL) + nccl4py ---- +# THE crux: tensorrt_llm's is_nccl_ep_installed() gates on libnccl >= 2.30.4, while the NGC +# TRT-LLM container bakes an older nvidia-nccl-cu13 — so NcclEP is dead on arrival as +# shipped. --no-deps so pip does not drag torch/TRT-LLM's pinned dependency graph backwards; +# we are deliberately overriding exactly one pin. nccl4py 0.3.1 ships the `nccl.ep` python +# package + libnccl_ep.so 0.1.0 (the version whose HT-kernel int64/int32 ABI detail the README +# documents). This nccl4py pin is REQUIRED, not merely current: nccl4py 0.4.1 ships no `nccl.ep` +# package at all (only nccl.bindings + nccl.core), so a bump to latest breaks `import nccl.ep`. +# Why 2.30.4 and not the newer 2.31.2: 2.31.2 also clears the >= 2.30.4 floor, but libnccl_ep 0.1.0 +# (shipped by the nccl4py pin above) is built against the NCCL 2.30 device-side API that the GIN +# CPU-proxy path calls into — and that device API is NOT append-only across 2.30.4->2.31.2. In the +# device struct libnccl_ep dereferences by pointer (ncclDevComm), 2.31.2 inserts hybridDenseGinBarrier +# at field 10 (before lsaMultimem) + backendIndex mid-GIN-block, and shrinks the by-value +# resourceWindow_inlined member (dropped its reserved padding) — each shifts the byte offsets of the +# GIN fields the CPU-proxy kernels read (verified by diffing nccl_device/impl/impl_comm__types.h at +# tags v2.30.4-1 vs v2.31.2-1). ncclGinType_t is append-only (PROXY=2 unchanged, +EFA_GDA=5), so the +# enum is not the issue — the devComm layout is. This is the silent-corruption class, so a build-only +# check cannot catch it. 2.30.4 is the version the 2026-08-07 correctness E2E actually ran. +# Two single-variable 2.31.2 trials on cgk p5en (--build-arg NVIDIA_NCCL_CU13=2.31.2) established: +# (+) LIBRARY ABI is clean on 2.31.2 — is_nccl_ep_installed() passes and the CommunicationFactory +# selects NcclEP on all 16 ranks (no quiet symbol/binding break). MEASURED. +# (=) DEVICE dispatch/combine round-trip: NOT YET measured on either arm. First trial's GIN init +# hard-failed because cgk's host gdrdrv is 2.4 (< the sample's GDRCopy-2.5 prereq); a second +# trial (2026-09-01) with a throwaway forced_pcie_copy instrument patch DID init GIN on +# gdrdrv-2.4 and both arms reached NcclEP selection, but crashed one line before the first +# dispatch on a latent probe-diagnostic bug (recipe/probe_nccl_ep.py referenced a +# NcclEpContext._ep_algorithm attr absent on rc24 — now getattr-guarded). See VERDICT.md. +# So the 2.31.2 device-side ABI is NOT yet certified (nor refuted) at the kernel level. Held at +# 2.30.4 as the measured-matching floor, not an upper bound — to bump it, re-run verify-image.sh + +# run-kernel-test.sh (probe fix now landed) on a gdrdrv>=2.5 host (or with the instrument patch) so +# GIN inits and the dispatch/combine round-trip + oracle actually execute on both arms. +ARG NVIDIA_NCCL_CU13=2.30.4 +ARG NCCL4PY_VER=0.3.1 +RUN pip3 install --no-cache-dir --no-deps --force-reinstall "nvidia-nccl-cu13==${NVIDIA_NCCL_CU13}" \ + && pip3 install --no-cache-dir "nccl4py==${NCCL4PY_VER}" \ + && NCCL_ROOT=$(python3 -c "import nvidia.nccl, pathlib; print(pathlib.Path(nvidia.nccl.__path__[0]))") \ + && ln -sf "$NCCL_ROOT/lib/libnccl.so.2" /usr/local/lib/libnccl.so.2 \ + && ln -sf "$NCCL_ROOT/lib/libnccl.so.2" /usr/local/lib/libnccl.so \ + && echo "$NCCL_ROOT/lib" > /etc/ld.so.conf.d/00-pip-nccl.conf && ldconfig \ + && [ "$(strings "$NCCL_ROOT/lib/libnccl.so.2" | grep -c "NCCL version ${NVIDIA_NCCL_CU13}")" -ge 1 ] +# The version assert uses the draining count form, not `grep -q` (SIGPIPE-141 flake under +# pipefail). Import checks (nccl.ep, is_nccl_ep_installed) live in recipe/verify-image.sh, +# which runs with GPUs — the build sandbox has none. + +# ---- Layer 5: aws-ofi-nccl GIN plugin (in-tree script, COPY'd not curled) ---- +# Built from the released tag v1.21.1 (same tag the sibling deepep-v2-benchmark builds): it +# exports the CPU-proxy GIN op-tables (ncclGinPlugin_v11/_v13) NCCL_GIN_TYPE=2 uses, and its +# forced-PCIe-with-fallback gdrcopy path is the released default — so no closed-PR cherry-pick +# and no OFI_NCCL_GDRCOPY_FORCED_PCIE_COPY override are needed. gdrdrv >= 2.5 is a host +# precondition (README Prerequisites) rather than a private plugin fork. +COPY setup_trtllm_nccl_ep_efa.sh /opt/setup_trtllm_nccl_ep_efa.sh +ARG AWS_OFI_NCCL_REF=v1.21.1 +RUN chmod +x /opt/setup_trtllm_nccl_ep_efa.sh \ + && AWS_OFI_NCCL_REF=${AWS_OFI_NCCL_REF} /opt/setup_trtllm_nccl_ep_efa.sh +ENV LD_LIBRARY_PATH=/opt/aws-ofi-nccl/lib:${LD_LIBRARY_PATH} +# NCCL_GIN_PLUGIN pairs with NCCL_NET_PLUGIN — one .so supplies both the net and GIN tables; +# set both as ENV (matching the sibling deepep-v2-benchmark) so the image is correct for +# anyone running it outside the launchers. NVIDIA_GDRCOPY=enabled matches the two sibling +# GIN images and the non-privileged/device-plugin path the manifest header offers. +ENV NCCL_NET_PLUGIN=/opt/aws-ofi-nccl/lib/libnccl-net-ofi.so \ + NCCL_GIN_PLUGIN=/opt/aws-ofi-nccl/lib/libnccl-net-ofi.so \ + NVIDIA_GDRCOPY=enabled + +# ---- Layer 6 (OPT-IN, default OFF): the HIGH_THROUGHPUT+FLAT selectability patch ---- +# Upstream's NcclEP hardcodes LOW_LATENCY + RANK_MAJOR. On THIS substrate (NCCL 2.30.4 + +# nccl_ep 0.1.0 + GIN CPU-proxy) the upstream default runs clean — measured both in a real +# serve and at 16 ranks cross-node — so the UNPATCHED image is the baseline and this layer +# defaults OFF. NVIDIA/TensorRT-LLM PR #17715 (open) makes algorithm/layout selectable via +# TRTLLM_NCCL_EP_ALGO / TRTLLM_NCCL_EP_LAYOUT; build with APPLY_HT_FLAT_PATCH=1 to bake the +# PR's three change commits (pinned at their immutable SHAs) into the container's site-packages. +# Fail-loud: if a hunk no longer applies against this base tag, the BUILD fails — do not +# ship an image whose patch state is ambiguous. Retire this layer when #17715 merges. +ARG APPLY_HT_FLAT_PATCH=0 +# The three single-parent commits of PR#17715 (NOT the branch-head merge commit +# e14a6f64 — GitHub's .patch endpoint 403s for a merge, so git format-patch emits nothing +# and the layer would die on it every time). +ARG HT_FLAT_PATCH_SHAS="6035d66353142d70ff41041b13dda1b3e788371a 2e3de8a3ccbb6bb1f78cd1504852045db989abfb 4e7cba789f4d928babc6199e895961f54ee30e11" +RUN set -euo pipefail; \ + if [ "${APPLY_HT_FLAT_PATCH}" = "1" ]; then \ + SP_PARENT=$(python3 -c "import tensorrt_llm, pathlib; print(pathlib.Path(tensorrt_llm.__file__).parent.parent)"); \ + cd "$SP_PARENT"; \ + for sha in ${HT_FLAT_PATCH_SHAS}; do \ + echo "== applying NVIDIA/TensorRT-LLM PR#17715 commit ${sha} =="; \ + curl -fsSL "https://github.com/NVIDIA/TensorRT-LLM/commit/${sha}.patch" > /tmp/${sha}.patch; \ + [ -s /tmp/${sha}.patch ] || { echo "FATAL: PR#17715 ${sha} yielded an empty patch (a merge commit has no format-patch output) — pin a single-parent commit"; exit 1; }; \ + git apply --include='tensorrt_llm/*' --check /tmp/${sha}.patch \ + || { echo "FATAL: PR#17715 ${sha} does not apply on ${TRTLLM_BASE} — base moved under the patch"; exit 1; }; \ + git apply --include='tensorrt_llm/*' /tmp/${sha}.patch; \ + rm -f /tmp/${sha}.patch; \ + done; \ + find "$SP_PARENT/tensorrt_llm/_torch/modules/fused_moe" -name '__pycache__' -prune -exec rm -rf {} + || true; \ + touch /opt/.ht-flat-patch-applied; \ + else echo "HT/FLAT patch layer skipped (APPLY_HT_FLAT_PATCH=0 — upstream LL/RANK_MAJOR default)"; fi + +# ---- Layer 7: the recipe scripts (LAST — script iteration never invalidates heavy layers) ---- +COPY recipe/serve.sh /opt/serve.sh +COPY recipe/run-kernel-test.sh /opt/run-kernel-test.sh +COPY recipe/probe_nccl_ep.py /opt/probe_nccl_ep.py +COPY recipe/benchmark_probe.py /opt/benchmark_probe.py +COPY recipe/benchmark.sh /opt/benchmark.sh +RUN chmod 755 /opt/serve.sh /opt/run-kernel-test.sh /opt/benchmark.sh + +CMD ["/bin/bash", "-lc", "echo 'run: /opt/serve.sh (single-node EP serve) | /opt/run-kernel-test.sh {leader|worker} (cross-node NcclEP probe)'; sleep infinity"] diff --git a/examples/inference/tensorrt-llm/nccl-ep-efa/README.md b/examples/inference/tensorrt-llm/nccl-ep-efa/README.md new file mode 100644 index 000000000..4347d1942 --- /dev/null +++ b/examples/inference/tensorrt-llm/nccl-ep-efa/README.md @@ -0,0 +1,224 @@ + +# TensorRT-LLM NcclEP — MoE expert-parallel all-to-all over AWS EFA (NCCL-GIN CPU-proxy) + +This test case runs **TensorRT-LLM's own MoE expert-parallel backend, `NcclEP`**, over **AWS EFA**. +Unlike the sibling DeepEP samples (vLLM, SGLang), TRT-LLM's wide-EP dispatch/combine path is not the +`deep_ep` Python package at all — it is **`nccl.ep`** (the `nccl4py` package + `libnccl_ep`), which +drives its network traffic **through NCCL**, and on EFA that means **aws-ofi-nccl with GIN +(GPU-Initiated Networking) compiled in**, running GIN's **CPU-proxy** mode. The image is built +NGC-from-scratch from public sources only; the recipe takes you from image build → transport smoke +test → a served `/v1/chat/completions` → a concurrency benchmark. + +## Pins (every one justified) + +| Component | Pin | Why | +|---|---|---| +| Base image | `nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc24` | GA v1.2.1 has **no NcclEP backend at all** (`nccl_ep_utils.py` absent at that tag), so a 1.3.0rc pin is *required*, not a preference; rc24 is the first release container whose factory ships the `NCCL_EP` arm natively. Re-test + re-pin when a GA carrying the backend appears. | +| EFA installer | 1.50.0 | The current EFA userspace and what the sibling deepep-v2-benchmark builds; ships aws-ofi-nccl 1.21.1 in-box. BUILD-PIN: the 2026-08-07 correctness E2E ran on 1.48.0, so this assembly is not yet cluster-re-measured on 1.50.0 (that is what `verify-image.sh` + `run-kernel-test.sh` exist to re-verify on-cluster). | +| NCCL | pip `nvidia-nccl-cu13==2.30.4` | `is_nccl_ep_installed()` gates on ≥ 2.30.4; the NGC container bakes an older NCCL, so NcclEP is dead-on-arrival as shipped (trap 3). | +| nccl4py | 0.3.1 (ships `nccl.ep` + `libnccl_ep` 0.1.0) | REQUIRED, not merely current: 0.4.1 ships no `nccl.ep` package at all (only `nccl.bindings`/`nccl.core`), so bumping to latest breaks `import nccl.ep`. It is `libnccl_ep` that is 0.1.0 (its HT-kernel int64 ABI detail matters — trap 4), not nccl4py. | +| aws-ofi-nccl | released tag `v1.21.1` | The released tag that carries the CPU-proxy GIN op-tables (`ncclGinPlugin_v11`/`_v13`) `NCCL_GIN_TYPE=2` uses, and the tag the sibling deepep-v2-benchmark builds from source. Its forced-PCIe-with-fallback gdrcopy path is the released default, so no closed-PR cherry-pick and no `OFI_NCCL_GDRCOPY_FORCED_PCIE_COPY` override are needed (the retired dev commit `9c44d34` + closed PR#1351 both dropped). | +| gdrcopy | commit `c91ad9f` (= v2.5.2) | GIN **requires** gdrcopy compiled in (trap 7). Commit pin, not tag. | +| GPU arch | validated on **H200 (p5en.48xlarge)** only | H100/p5 differs only in EFA NIC count (32 vs 16) in the manifest; not re-measured there. | + +## How NcclEP gets onto EFA + +``` +trtllm-serve ── CommunicationFactory.create_strategy() (reads TRTLLM_FORCE_COMM_METHOD) + └─ NcclEP ── nccl.ep (nccl4py / libnccl_ep 0.1.0) + └─ NCCL 2.30.4 (pip) ── GIN device API (NCCL_GIN_TYPE=2, CPU-proxy) + └─ aws-ofi-nccl @v1.21.1 (ncclGinPlugin, gdrcopy) + └─ libfabric ── EFA (efa-direct, SRD) +``` + +Selection is *forced* with `TRTLLM_FORCE_COMM_METHOD=NCCL_EP`, but only after TRT-LLM's own +feasibility gate (`_get_nccl_ep_unavailable_reason`) passes and **only when attention DP is +enabled** (trap 2). The proof that the whole chain engaged is one log line, on every rank: + +``` +[TRT-LLM] [RANK 5] [I] NCCL EP group created: ep_size=8, num_experts=128, ..., + layout=RANK_MAJOR, algorithm=LOW_LATENCY +``` + +## The integration traps + +1. **`TLLM_LOG_LEVEL=info`, not `TRTLLM_LOG_LEVEL`.** `tensorrt_llm/logger.py` reads `TLLM_LOG_LEVEL`; + the intuitive `TRTLLM_`-prefixed spelling is **silently ignored**, the level stays at `error`, and + the one "NCCL EP group created" line that constitutes engagement proof never prints. You then have + a serve that answers requests while you cannot tell which all-to-all it is running. +2. **No attention DP ⇒ the env knob is never read.** `create_strategy()` returns `None` / + `AllGatherReduceScatter` when `enable_attention_dp` is false or `dp_size == 1` — *before* it looks + at `TRTLLM_FORCE_COMM_METHOD`. `recipe/serve.sh` passes `enable_attention_dp: true` via + `--extra_llm_api_options`; without it the serve comes up fine on the wrong backend. +3. **The NGC container's baked NCCL is too old for NcclEP.** `is_nccl_ep_installed()` requires + NCCL ≥ 2.30.4. The Dockerfile force-installs pip `nvidia-nccl-cu13==2.30.4` (`--no-deps`, one pin + overridden deliberately) and every launcher puts the pip wheel's `lib/` **first** on + `LD_LIBRARY_PATH` — if the baked copy wins the link, NcclEP silently reports unavailable. + `recipe/verify-image.sh` asserts which `libnccl.so.2` resolves *and* its version string. +4. **Harness-green ≠ serve-green (the int64/int32 boundary).** The nccl_ep 0.1.x HIGH_THROUGHPUT + kernel ABI hands back `recv_topk_idx` as **int64**, while the downstream consumer + `torch.ops.trtllm.fused_moe` requires **int32** routing ids — a real model forward crashes + ("token_selected_experts dtype is Long") even though a standalone dispatch/combine harness passes + 16/16, because the harness does its own expert math and never crosses that boundary. Under the + upstream LOW_LATENCY/RANK_MAJOR default the buffer is already int32, so the conflict is invisible. + [TensorRT-LLM PR#17715](https://github.com/NVIDIA/TensorRT-LLM/pull/17715) carries the + dispatch-return narrowing with the selectability patch (Dockerfile `APPLY_HT_FLAT_PATCH=1` layer). + There is also a **second, package-side exit** the code already anticipates: TRT-LLM gates on + `_MIN_NCCL_EP_INT32_TOPK_VERSION = "0.2"`, so a `libnccl_ep >= 0.2` (we pin 0.1.0) would return + int32 topk ids natively and retire this boundary with **no patch at all**. It is unreachable today + — no such nccl4py wheel is published — so it is a PyPI watch to track next to PR#17715, not a + current option; whichever lands first makes the opt-in patch layer unnecessary. +5. **The GA-that-lacks-the-backend trap.** Pinning "the latest GA" (v1.2.1) gives you an image where + this entire test case is impossible — the factory modules do not exist at that tag. + `recipe/verify-image.sh` fails loud on that with "wrong base tag?". +6. **Do not carry "LOW_LATENCY faults on EFA" as a rule.** On *this* substrate (NCCL 2.30.4 + + nccl_ep 0.1.0 + GIN CPU-proxy) the upstream LL/RANK_MAJOR default was measured clean — a real EP8 + serve answered correctly and a 16-rank cross-node probe passed with zero illegal-memory-access. + An LL fault seen elsewhere was a different substrate. The HT/FLAT patch layer is therefore an + **opt-in selector** (default OFF), not a fix for a broken default. +7. **GIN needs gdrcopy at compile time and gdrdrv at run time.** An aws-ofi-nccl built without + `gdrapi.h` carries "GDRCopy support not available at compile time" and GIN init fails at serve; + the setup script asserts that string is *absent* from the built plugin. At run time, clusters + without a gdrdrv device plugin need `privileged: true` (the unprivileged device cgroup blocks + `open("/dev/gdrdrv")` with EPERM even after CAP_MKNOD — measured; the manifest header documents it). +8. **A node-spanning `trtllm-serve` does not bootstrap on EKS.** trtllm-serve uses mpirun across + nodes, and mpirun's OOB cannot route between VPC-CNI pods (each pod has a /32 eth0, so OMPI's + `opal_net_samenetwork()` never matches a peer). torchrun is unaffected — which is why this sample's + cross-node proof is the torchrun probe and the served proof is single-node EP8 (Known limitations). + +## Runtime requirements + +Two groups, and they live in different places by design. The **NCCL-GIN + EFA transport contract** +is identical on both pods, so it is set in the manifest env block *and* re-exported by every +launcher (the manifest alone documents the transport). The **NcclEP-selection knobs** are +role-specific — the serve forces them, the idle probe-peer must not — so the launchers +(`serve.sh` / `run-kernel-test.sh`) own them and they are deliberately absent from the manifest env +(a `TRTLLM_FORCE_COMM_METHOD` on the idle pod would be misleading; the probe launcher sets it when +that pod actually joins the all-to-all). + +Transport contract (manifest env + every launcher): + +| Env | Why | +|---|---| +| `NCCL_GIN_TYPE=2`, `NCCL_GIN_ENABLE=1` | GIN CPU-proxy — the EFA-viable GIN mode | +| `FI_PROVIDER=efa`, `FI_EFA_USE_DEVICE_RDMA=1` | EFA with GPU-direct RDMA | +| `FI_EFA_ENABLE_SHM_TRANSFER=0`, `FI_EFA_FORK_SAFE=1` | no SHM shortcut; fork-safe for the proxy | +| `NCCL_NET_PLUGIN=/opt/aws-ofi-nccl/lib/libnccl-net-ofi.so` | the GIN-capable plugin, explicitly | +| `NCCL_CUMEM_ENABLE=1`, `NCCL_NVLS_ENABLE=0`, `NCCL_IGNORE_DISABLED_P2P=1` | NCCL settings the measured substrate ran | +| `NCCL_NET_PLUGIN` = `NCCL_GIN_PLUGIN` = the built plugin `.so` | one plugin supplies both the net and GIN tables (also baked as image ENV) | + +NcclEP-selection knobs (launcher-owned — `serve.sh` and `run-kernel-test.sh`, set identically so the +correctness probe exercises the served config; NOT in the manifest env): + +| Env | Why | +|---|---| +| `TRTLLM_FORCE_COMM_METHOD=NCCL_EP` | selects NcclEP in the factory (trap 2 applies) | +| `NCCL_EP_NUM_QP_PER_RANK=32` | QP fan-out per rank on EFA — throughput-relevant (benchmark provenance) | +| `ENABLE_CONFIGURABLE_MOE=1` | exposes the configurable-MoE path NcclEP routes through | +| `TLLM_LOG_LEVEL=info` | prints the one "NCCL EP group created" engagement line (trap 1 — NOT `TRTLLM_LOG_LEVEL`) | + +## Prerequisites + +- An EKS cluster with p5en.48xlarge (or p5.48xlarge — adjust the EFA count in the manifest) nodes, + the EFA device plugin, the NVIDIA device plugin, and 2Mi hugepages pre-allocated on the nodes. +- The **`gdrdrv` kernel module loaded on the host** (`lsmod | grep gdrdrv` must be non-empty, so + `/dev/gdrdrv` exists). GIN needs GDRCopy at run time (trap 7); the manifest's `privileged: true` + is what lets the container open the host's `/dev/gdrdrv`, but privileged cannot conjure the device + node if the module was never loaded. The AWS GPU AMIs ship it; if absent, `sudo modprobe gdrdrv` + (gdrcopy >= 2.5, matching the image's `c91ad9f`/v2.5.2 build). +- Docker + network access to `nvcr.io`, `pypi.org`, `github.com`, `efa-installer.amazonaws.com`. +- An image registry you own (ECR); **do not point at anyone's private registry**. + +## Build + +```bash +cp setup/env_vars.example setup/env_vars # edit REGISTRY etc. (env_vars is gitignored) +setup/build-push.sh # baseline image (upstream LL/RANK_MAJOR default) +APPLY_HT_FLAT_PATCH=1 setup/build-push.sh # opt-in: bake PR#17715 (algorithm/layout selectable) +``` + +Then gate the image before any model load. `build-push.sh` sources `setup/env_vars` in its own +child shell, so `REGISTRY`/`IMAGE_NAME`/`IMAGE_TAG` do not reach your shell — source it here too, +and reference the same `${IMAGE_NAME}:${IMAGE_TAG}` the build used so the two cannot drift: + +```bash +source setup/env_vars +recipe/verify-image.sh ${REGISTRY}/${IMAGE_NAME}:${IMAGE_TAG} +# ... ALL CHECKS PASS +``` + +## Smoke-test the transport before the model load + +Deploy `kubernetes/trtllm-nccl-ep-2node.yaml` (set `image:` to your pushed image), then run the +16-rank cross-node probe — it drives TRT-LLM's real factory entrypoint and checks dispatch/combine +numerics against an oracle, in minutes, before any weights download: + +```bash +# LEADER is the stable headless-Service DNS name — it survives pod restarts (a raw pod IP +# does not) and needs no lookup; this is what the headless Service + publishNotReadyAddresses +# in the manifest exist to provide. +LEADER=trtllm-nccl-ep-0.trtllm-nccl-ep.trtllm-nccl-ep.svc.cluster.local +kubectl -n trtllm-nccl-ep exec trtllm-nccl-ep-1 -- bash -lc \ + "nohup /opt/run-kernel-test.sh worker $LEADER 1 > /tmp/kt.log 2>&1 &" +kubectl -n trtllm-nccl-ep exec trtllm-nccl-ep-0 -- /opt/run-kernel-test.sh leader $LEADER +# ... PROBE-PASS factory-selected NcclEP LOW_LATENCY+RANK_MAJOR world=16 +# ... KERNEL-TEST PASS — factory-selected NcclEP dispatch/combine over EFA verified +``` + +## Serve + +Pod 0 of the manifest runs `recipe/serve.sh` (single-node, `tp_size=ep_size=8`, +Qwen/Qwen3-30B-A3B by default — any MoE whose routed-expert count divides by `SERVE_EP` passes the +preflight). Confirm engagement, then ask it something with a checkable answer: + +```bash +kubectl -n trtllm-nccl-ep logs trtllm-nccl-ep-0 | grep "NCCL EP group created" +# curl the serve from inside pod 0 (127.0.0.1 = the serve's own /health address); from your +# workstation instead, `kubectl -n trtllm-nccl-ep port-forward svc/trtllm-nccl-ep 8000:8000` +# and curl 127.0.0.1:8000 — either way no raw pod IP, which a restart would change. +kubectl -n trtllm-nccl-ep exec trtllm-nccl-ep-0 -- \ + curl -s http://127.0.0.1:8000/v1/chat/completions -H 'Content-Type: application/json' -d \ + '{"model":"Qwen/Qwen3-30B-A3B","messages":[{"role":"user","content":"What is 17 multiplied by 23? Reply with only the number."}],"max_tokens":16}' +# a correct "391" pins routing correctness, not mere fluency +``` + +## Benchmark + +```bash +kubectl -n trtllm-nccl-ep exec trtllm-nccl-ep-0 -- \ + bash -lc 'OUT_ROOT=/work/benchmarks /opt/benchmark.sh 127.0.0.1' +``` + +Methodology, provenance knobs, and status live in [benchmarks/README.md](benchmarks/README.md). + +## Known limitations + +- **The served completion is single-node (EP8).** The cross-node (EP16) proof is at the + factory/dispatch level via the torchrun probe, not a served HTTP completion — a node-spanning + `trtllm-serve` needs mpirun, which cannot bootstrap between VPC-CNI pods (trap 8). Both halves + were measured; the combined "one serve spanning nodes" was not, and this sample does not claim it. +- **This exact image assembly is build-staged, not yet cluster-re-measured.** The measured E2E + (2026-08-07: real `trtllm-serve` HTTP-200-correct EP8 + 16-rank cross-node probe 16/16 PASS, + zero illegal-memory-access, `efa-direct` on every rank, both LL/RANK_MAJOR and patched HT/FLAT) + ran the same component versions with the rc24 NcclEP modules grafted onto an rc9 container. This + Dockerfile builds the cleaner equivalent — rc24 native — which is exactly what the recipe's gates + (`verify-image.sh`, `run-kernel-test.sh`) exist to re-verify on your cluster. +- **No performance numbers are published here** (see benchmarks/README.md). Nothing in this sample + demonstrates a speedup; it demonstrates a *working, verifiable* NcclEP-over-EFA path. +- Validated on H200/p5en only; p5/H100 differs in manifest EFA count and was not re-measured. + +## References + +- [NVIDIA/TensorRT-LLM PR#17715](https://github.com/NVIDIA/TensorRT-LLM/pull/17715) — NcclEP + algorithm/layout selectability (`TRTLLM_NCCL_EP_ALGO` / `TRTLLM_NCCL_EP_LAYOUT`) + the int64→int32 + dispatch-return fix; the Dockerfile's opt-in patch layer, retired when it merges. +- [NVIDIA/TensorRT-LLM issue#17714](https://github.com/NVIDIA/TensorRT-LLM/issues/17714) — the + companion issue documenting the HT/FLAT selection gap. +- [aws/aws-ofi-nccl](https://github.com/aws/aws-ofi-nccl) tag `v1.21.1` — the GIN plugin + source. Its `gdr_pin_buffer_v2` already attempts `GDR_PIN_FLAG_FORCE_PCIE` and falls back on + failure, so the released default covers what the retired dev pin's `OFI_NCCL_GDRCOPY_FORCED_PCIE_COPY` + override (closed PR#1351) used to force — no cherry-pick needed. +- Sibling test cases: [SGLang + DeepEP over EFA](../../sglang/dsr1-deepep-efa/) (NVSHMEM host-proxy + substrate), the [expert-parallelism micro-benchmarks](../../../../micro-benchmarks/expert-parallelism/), + and — once [#1230](https://github.com/awslabs/awsome-distributed-ai/pull/1230) merges — vLLM + + DeepEP-V2 over EFA (same NCCL-GIN CPU-proxy substrate, different EP kernel package). diff --git a/examples/inference/tensorrt-llm/nccl-ep-efa/benchmarks/README.md b/examples/inference/tensorrt-llm/nccl-ep-efa/benchmarks/README.md new file mode 100644 index 000000000..ccf5f0409 --- /dev/null +++ b/examples/inference/tensorrt-llm/nccl-ep-efa/benchmarks/README.md @@ -0,0 +1,59 @@ + +# Benchmarks — TensorRT-LLM NcclEP over EFA + +## Status: correctness measured, performance NOT yet published + +What **was** measured (2026-08-07, 2× p5en.48xlarge / H200, same component versions as this image +with the rc24 NcclEP modules grafted onto an rc9 container — see the sample README "Known +limitations"): + +- A real `trtllm-serve` (Qwen/Qwen3-30B-A3B, single-node EP8) reached `Application startup + complete`, logged `NCCL EP group created` on all 8 ranks, and answered `/v1/chat/completions` + correctly (including an arithmetic check — MoE misrouting produces plausible text but wrong + arithmetic). +- The 16-rank cross-node probe (`recipe/run-kernel-test.sh`) passed 16/16 with zero + illegal-memory-access and the `efa-direct` provider banner on every rank, under **both** the + upstream LOW_LATENCY/RANK_MAJOR default and (on a patched image) HIGH_THROUGHPUT/FLAT, with + combine relmax ≈ 0.003–0.005 against the oracle. + +What was **not** measured, and is therefore not claimed here: + +- **No throughput/latency tables are published in this folder yet.** Run `recipe/benchmark.sh` + on your own deployment; treat a single pass as directional. +- **No alternative-backend baseline was measured** — nothing here compares NcclEP against + TRT-LLM's default all-to-all or any other EP backend, so no relative-performance claim is made. +- No p5/H100 numbers; no multi-node *served* numbers (single-node serve boundary — README trap 8). + +## Methodology (what `benchmark.sh` runs and why) + +`recipe/benchmark_probe.py` sweeps fixed concurrency levels against the live OpenAI endpoint: + +- **Requests per level = concurrency × 5** so p50/p90/p99 describe a distribution, not one shot. +- **Successes only** enter the percentiles and the token numerator; any failure fails the whole + run (exit ≠ 0) — a partially-dead serve must not report a *better* p50 because refused + connections return fast. +- **`ignore_eos: true`** pins generated tokens == `max_tokens`, fixing the tok/s denominator by + construction. +- **Unique prompt prefix per request** (the index goes first) so a prefix/KV cache cannot serve + request N's prefill from request 1's. +- A 200 response **without a `usage` block is a failure** (`200-no-usage`), not a zero-token + success that silently deflates throughput. +- The probe's **exit code is the sole pass/fail authority** — `benchmark.sh` adds no separate + grep gate that could disagree with it. +- Default levels are `[1, 2, 4]`, sized to the measured serve shape (`--max_batch_size 4`); + raise `SERVE_MAX_BATCH_SIZE` / `SERVE_MAX_NUM_TOKENS` in `serve.sh` before sweeping higher, + and record both. + +## Provenance knobs (record these next to any numbers you publish) + +| Knob | Default | Why it belongs in the table | +|---|---|---| +| `TRTLLM_NCCL_EP_ALGO` (+ layout) | `LOW_LATENCY` (+RANK_MAJOR) | different EP kernels entirely; patched image required for non-default (README trap 4) | +| `NCCL_EP_NUM_QP_PER_RANK` | 32 | QP fan-out per rank on EFA — throughput-relevant | +| `SERVE_TP` / `SERVE_EP` | 8 / 8 | the parallel shape | +| `SERVE_MAX_BATCH_SIZE` / `SERVE_MAX_NUM_TOKENS` / `SERVE_MAX_SEQ_LEN` | 4 / 2048 / 2048 | serving-engine ceilings that bound every level | +| model | `Qwen/Qwen3-30B-A3B` | expert count/hidden size change the all-to-all shape | +| instance type / EFA NICs | p5en.48xlarge / 16 | fabric budget per node | + +Raw sweeps land in `benchmarks/raw//` (gitignored — publish tables here, keep raw +JSONL out of the repo). diff --git a/examples/inference/tensorrt-llm/nccl-ep-efa/kubernetes/trtllm-nccl-ep-2node.yaml b/examples/inference/tensorrt-llm/nccl-ep-efa/kubernetes/trtllm-nccl-ep-2node.yaml new file mode 100644 index 000000000..2468394e5 --- /dev/null +++ b/examples/inference/tensorrt-llm/nccl-ep-efa/kubernetes/trtllm-nccl-ep-2node.yaml @@ -0,0 +1,204 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. SPDX-License-Identifier: MIT-0 +# TensorRT-LLM NcclEP MoE all-to-all over AWS EFA — 2× p5en.48xlarge (H200). +# +# TOPOLOGY (honest shape — read before scaling expectations): +# ordinal 0: single-node trtllm-serve, tp_size=ep_size=8 (serve.sh) — the SERVED proof. +# ordinal 1: probe peer — idles, then joins the 16-rank CROSS-NODE NcclEP probe +# (run-kernel-test.sh) that proves dispatch/combine crosses the EFA node +# boundary. +# Why not one serve spanning both nodes: a node-spanning trtllm-serve needs mpirun, +# and mpirun's OOB cannot route between EKS pods (AWS VPC-CNI gives each pod a /32 +# eth0, so OMPI's opal_net_samenetwork() never matches a peer and the HNP reports +# "no route to" the remote orted). The torchrun-driven probe does not share that +# limitation — hence serve-on-one-node + probe-across-two. See README "Known limitations". +# +# Cross-node probe, once both pods are Running. LEADER is the stable headless-Service DNS +# name (this is what publishNotReadyAddresses + the headless clusterIP below are FOR — +# it survives pod restarts, where a raw pod IP does not, and needs no manual lookup): +# LEADER=trtllm-nccl-ep-0.trtllm-nccl-ep.trtllm-nccl-ep.svc.cluster.local +# kubectl -n trtllm-nccl-ep exec trtllm-nccl-ep-1 -- bash -lc \ +# "nohup /opt/run-kernel-test.sh worker $LEADER 1 > /tmp/kt.log 2>&1 &" +# kubectl -n trtllm-nccl-ep exec trtllm-nccl-ep-0 -- /opt/run-kernel-test.sh leader $LEADER +# Free ordinal 0's GPUs for the probe by deploying with RUN_SERVE=0 (env below) so it holds +# idle, then flip RUN_SERVE=1 to serve; or, if it is already serving, pkill -f '[t]rtllm-serve'. +# +# Scale seams (per-node EFA NIC count differs by instance type): +# p5.48xlarge (H100): vpc.amazonaws.com/efa: "32", nodeSelector p5.48xlarge +# p5en.48xlarge(H200): vpc.amazonaws.com/efa: "16" ← this file (the measured topology) +# +# Image: build the sample-root Dockerfile (NGC-from-scratch, public sources only) and push +# to YOUR registry — setup/build-push.sh builds ${REGISTRY}/${IMAGE_NAME}:${IMAGE_TAG} from +# setup/env_vars; set image: below to that SAME name. Do NOT point at anyone's private registry. +# +# privileged: true is REQUIRED where /dev/gdrdrv is not advertised by a device plugin +# (measured: the unprivileged device cgroup blocks open("/dev/gdrdrv") with EPERM even +# after CAP_MKNOD; without GDRCopy aws-ofi-nccl reports ginType=NONE and the NCCL-GIN +# CPU-proxy path is dead). If your cluster runs a gdrdrv device plugin, drop privileged +# and request the device instead. +--- +apiVersion: v1 +kind: Namespace +metadata: + name: trtllm-nccl-ep +--- +apiVersion: v1 +kind: Service +metadata: + name: trtllm-nccl-ep + namespace: trtllm-nccl-ep +spec: + clusterIP: None # headless — stable per-pod DNS for the probe rendezvous + # publishNotReadyAddresses pairs with the probes below: the probe peer resolves pod 0 + # through this Service's DNS A-record BEFORE the serve is up, so the not-ready address + # must still publish or the torchrun rendezvous deadlocks. + publishNotReadyAddresses: true + selector: + app: trtllm-nccl-ep + ports: + - { port: 8000, name: http } + - { port: 29501, name: probe-rdzv } +--- +apiVersion: apps/v1 +kind: StatefulSet +metadata: + name: trtllm-nccl-ep + namespace: trtllm-nccl-ep +spec: + serviceName: trtllm-nccl-ep + replicas: 2 # pod 0 = EP8 serve; pod 1 = cross-node probe peer (header) + podManagementPolicy: Parallel + selector: + matchLabels: + app: trtllm-nccl-ep + template: + metadata: + labels: + app: trtllm-nccl-ep + spec: + # one pod per node — EFA/GPU are node resources; same-node EFA loopback is + # silently dropped by SRD anyway, so co-scheduling two pods is never right + affinity: + podAntiAffinity: + requiredDuringSchedulingIgnoredDuringExecution: + - labelSelector: + matchLabels: + app: trtllm-nccl-ep + topologyKey: kubernetes.io/hostname + nodeSelector: + node.kubernetes.io/instance-type: p5en.48xlarge + tolerations: + - { key: nvidia.com/gpu, operator: Exists, effect: NoSchedule } + terminationGracePeriodSeconds: 30 + containers: + - name: trtllm + image: REPLACE_WITH_YOUR_REGISTRY/trtllm-nccl-ep-efa:v1-20260825 # = ${REGISTRY}/${IMAGE_NAME}:${IMAGE_TAG} from setup/env_vars + imagePullPolicy: IfNotPresent + command: ["/bin/bash", "-c"] + args: + - | + set -euo pipefail + echo "=== trtllm-nccl-ep pod $(hostname) $(date -u +%FT%TZ) ===" + echo "aws-ofi-nccl effective SHA: $(cat /opt/aws-ofi-nccl.effective.sha 2>/dev/null || echo unknown)" + ORDINAL=${HOSTNAME##*-} + # RUN_SERVE gates whether ordinal 0 becomes the serve (PID 1). Deploy once with + # RUN_SERVE=0 to hold ordinal 0 idle for the cross-node probe (the header's + # deploy->probe->serve order — otherwise serve is PID 1 and the "pkill -f + # [t]rtllm-serve" step has nothing to hand the GPUs to), then flip to 1 to serve. + if [ "$ORDINAL" = "0" ] && [ "${RUN_SERVE:-1}" = "1" ]; then + exec /opt/serve.sh + else + echo "probe peer ready — join the cross-node kernel test with:" + echo " /opt/run-kernel-test.sh worker trtllm-nccl-ep-0.trtllm-nccl-ep.trtllm-nccl-ep.svc.cluster.local $ORDINAL" + touch /tmp/probe-peer.ready + exec sleep infinity + fi + env: + # ---- model + shape knobs (defaults = the measured EP8 serve) ---- + # RUN_SERVE=1 (default) makes ordinal 0 the EP8 serve; set 0 to hold it idle as a + # probe peer for the deploy->probe->serve order (args block + header). The probes + # below key off it too, so RUN_SERVE=0 does not restart-loop waiting on :8000/health. + - { name: RUN_SERVE, value: "1" } + - { name: SERVE_MODEL, value: "Qwen/Qwen3-30B-A3B" } # any MoE with routed-experts % EP == 0 (serve.sh asserts) + - { name: SERVE_TP, value: "8" } + - { name: SERVE_EP, value: "8" } + - { name: TRTLLM_NCCL_EP_ALGO, value: "LOW_LATENCY" } # HIGH_THROUGHPUT needs APPLY_HT_FLAT_PATCH=1 image (serve.sh refuses otherwise) + # - { name: HF_TOKEN, valueFrom: { secretKeyRef: { name: hf-token, key: token } } } # gated models only + + # ---- NCCL-GIN proxy + EFA contract, VERBATIM (also exported by serve.sh and + # run-kernel-test.sh; repeated here so the manifest alone documents the + # transport contract) ---- + - { name: NCCL_GIN_TYPE, value: "2" } # 2 = CPU-proxy GIN (the EFA-viable path) + - { name: NCCL_GIN_ENABLE, value: "1" } + - { name: NCCL_CUMEM_ENABLE, value: "1" } + - { name: NCCL_NVLS_ENABLE, value: "0" } + - { name: NCCL_IGNORE_DISABLED_P2P, value: "1" } + - { name: FI_PROVIDER, value: "efa" } + - { name: FI_EFA_USE_DEVICE_RDMA, value: "1" } + - { name: FI_EFA_ENABLE_SHM_TRANSFER, value: "0" } + - { name: FI_EFA_FORK_SAFE, value: "1" } + - { name: OFI_NCCL_PROTOCOL, value: "RDMA" } + - { name: NCCL_NET_PLUGIN, value: "/opt/aws-ofi-nccl/lib/libnccl-net-ofi.so" } + - { name: NCCL_DEBUG, value: "WARN" } # INFO to see the efa-direct banner proof + # pod 0 WHEN SERVING (RUN_SERVE=1): trtllm-serve /health on :8000 = the real "serving" + # signal (weights loaded, NcclEP groups created). Every other case — pod 1, or pod 0 + # held idle with RUN_SERVE=0 — is a probe peer, ready once its shell touches the file + # (so RUN_SERVE=0 does not restart-loop against a :8000 that will never come up). + startupProbe: + exec: + command: ["/bin/bash", "-c", "if [ \"${HOSTNAME##*-}\" = \"0\" ] && [ \"${RUN_SERVE:-1}\" = \"1\" ]; then curl -sf http://127.0.0.1:8000/health; else test -f /tmp/probe-peer.ready; fi"] + periodSeconds: 30 + failureThreshold: 60 # up to 30 min: 30B-class weight download + load + readinessProbe: + exec: + command: ["/bin/bash", "-c", "if [ \"${HOSTNAME##*-}\" = \"0\" ] && [ \"${RUN_SERVE:-1}\" = \"1\" ]; then curl -sf http://127.0.0.1:8000/health; else test -f /tmp/probe-peer.ready; fi"] + periodSeconds: 15 + # No livenessProbe by deliberate choice: a serve that wedges after startup drops out of + # the Service endpoints (readiness) but is NOT auto-restarted, so its logs/GPU state + # survive for inspection — the right trade for a two-pod diagnostic sample over a + # self-healing production serve. Add a livenessProbe (same command) to flip that. + # requests == limits (Guaranteed QoS): NCCL_GIN_TYPE=2 is the CPU-PROXY path, so + # proxy-thread CPU sits on the data path — Burstable QoS + CFS throttling under node + # pressure would degrade network throughput and silently taint published numbers. + resources: + limits: + nvidia.com/gpu: "8" + vpc.amazonaws.com/efa: "16" # p5en=16, p5=32 + hugepages-2Mi: 5120Mi # needs 2Mi hugepages pre-allocated on the node (they were on the nodes the measured runs used; without them the pod sits Pending) + cpu: "90" + memory: 1024Gi + ephemeral-storage: 200Gi # reserve the `work` emptyDir (below): sizeLimit alone bounds but does not RESERVE it, so the scheduler could place this on a node without room and evict mid-run + requests: + nvidia.com/gpu: "8" + vpc.amazonaws.com/efa: "16" + hugepages-2Mi: 5120Mi + cpu: "90" + memory: 1024Gi + ephemeral-storage: 200Gi + securityContext: + privileged: true # gdrdrv device-cgroup (see header) + capabilities: + # IPC_LOCK is redundant UNDER privileged (privileged already grants the full + # capability set) — kept so it stays load-bearing for the non-privileged/ + # device-plugin variant the header offers: drop `privileged`, add `--device + # /dev/gdrdrv`, and IPC_LOCK is then the capability RDMA memory pinning needs. + add: [IPC_LOCK] + ports: + - { containerPort: 8000, name: http } + - { containerPort: 29501, name: probe-rdzv } + volumeMounts: + - { name: shm, mountPath: /dev/shm } + - { name: dev-infiniband, mountPath: /dev/infiniband } + # NON-PRIVILEGED variant (drop `privileged: true` above): also mount /dev/gdrdrv, the + # way the sibling megatron-bridge and sglang samples do — GIN's gdrcopy handle needs + # the device node, which privileged grants implicitly but an unprivileged pod does not. + # - { name: dev-gdrdrv, mountPath: /dev/gdrdrv } + - { name: work, mountPath: /work } # HF cache + logs (per-pod, ephemeral) + volumes: + - name: shm + emptyDir: { medium: Memory, sizeLimit: 64Gi } + - name: dev-infiniband + hostPath: { path: /dev/infiniband } + # - name: dev-gdrdrv # pairs with the non-privileged mount above + # hostPath: { path: /dev/gdrdrv, type: CharDevice } + - name: work + emptyDir: { sizeLimit: 200Gi } # Qwen3-30B ~60GB of weights + HF cache; size for your model diff --git a/examples/inference/tensorrt-llm/nccl-ep-efa/recipe/benchmark.sh b/examples/inference/tensorrt-llm/nccl-ep-efa/recipe/benchmark.sh new file mode 100755 index 000000000..c796431f8 --- /dev/null +++ b/examples/inference/tensorrt-llm/nccl-ep-efa/recipe/benchmark.sh @@ -0,0 +1,14 @@ +#!/usr/bin/env bash +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. SPDX-License-Identifier: MIT-0 +# Concurrency sweep against a live serve. Writes fresh benchmarks/raw//; exits non-zero on any failure. +# The pass/fail authority is benchmark_probe.py's exit code (0 only if EVERY request at +# EVERY level succeeded) — no separate grep gate that could disagree with it. +set -euo pipefail +SERVE_IP="${1:?usage: benchmark.sh }" +# OUT_ROOT: benchmarks/raw in the repo checkout; inside the pod set OUT_ROOT=/work/benchmarks +OUT_DIR="${OUT_ROOT:-benchmarks/raw}/$(date -u +%Y%m%dT%H%M%SZ)"; mkdir -p "$OUT_DIR" +rc=0 +python3 "$(dirname "$0")/benchmark_probe.py" --url "http://${SERVE_IP}:8000/v1/chat/completions" \ + --out "${OUT_DIR}/sweep.jsonl" | tee "${OUT_DIR}/sweep.txt" || rc=$? +[ "$rc" -eq 0 ] || { echo "FAIL: probe reported failed requests (rc=$rc) — not a valid run"; exit "$rc"; } +echo "results in ${OUT_DIR}" diff --git a/examples/inference/tensorrt-llm/nccl-ep-efa/recipe/benchmark_probe.py b/examples/inference/tensorrt-llm/nccl-ep-efa/recipe/benchmark_probe.py new file mode 100644 index 000000000..545fc92d4 --- /dev/null +++ b/examples/inference/tensorrt-llm/nccl-ep-efa/recipe/benchmark_probe.py @@ -0,0 +1,150 @@ +#!/usr/bin/env python3 +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. SPDX-License-Identifier: MIT-0 +# Throughput + latency-vs-concurrency probe for the live trtllm-serve NcclEP EP serve. +# Stdlib only (urllib + concurrent.futures) so it runs in the pod with no pip. +# Methodology (fail-loud, distribution-honest): +# - each level fires requests-per-level = conc * NREQ_MULT requests (default 5x) at a +# fixed concurrency, so percentiles describe a distribution, not a single shot +# - only SUCCESSFUL requests enter the latency percentiles and the token numerator; +# failures are counted separately and fail the run (a partially-dead serve must not +# report a *better* p50 because refused connections return fast) +# - "ignore_eos": true pins every request to exactly max_tokens generated tokens, so +# the tok/s denominator is fixed by construction rather than model/prompt-dependent +# - each request carries a unique prompt PREFIX (the index goes first), so a prefix/KV +# cache cannot serve request N's prefill from request 1's +# - a missing usage block in a 200 response is a FAILURE (code "200-no-usage"), not a +# zero-token success that silently deflates throughput +# - exit 0 only if EVERY request at EVERY level succeeded +import argparse, json, os, time, urllib.request, urllib.error, sys +from concurrent.futures import ThreadPoolExecutor, as_completed + +# Defaults keep the probe runnable standalone inside the pod (127.0.0.1 loopback); +# recipe/benchmark.sh overrides --url (leader IP) and --out (raw JSONL path). +URL = "http://127.0.0.1:8000/v1/chat/completions" +MODEL = os.environ.get("SERVE_MODEL", "Qwen/Qwen3-30B-A3B") +MAX_TOKENS = 128 +PROMPT = "Explain expert parallelism in mixture-of-experts models in two sentences." +CONCURRENCIES = [int(c) for c in os.environ.get("BENCH_CONCURRENCIES", "1,2,4").split(",")] + # sized to the measured serve shape (--max_batch_size 4); raise + # SERVE_MAX_BATCH_SIZE/SERVE_MAX_NUM_TOKENS before sweeping higher. + # Env-read (not a module constant) so a higher sweep needs no image + # rebuild — the file is baked at /opt/benchmark_probe.py. +NREQ_MULT = 5 # requests per level = conc * this (>=5x so p50/p90 are distributions) +WARMUP = 3 # discarded requests before the FIRST measured level: level 0 otherwise absorbs + # lazy CUDA module load / allocator growth / tokenizer init as ~20% of its n=5 + # sample, inflating its percentiles and making the sweep look superlinear + +def one_request(url, idx): + body = json.dumps({ + "model": MODEL, + # unique prefix per request (index FIRST) so a prefix cache cannot make + # prefill free for requests 2..N — see benchmarks/README methodology notes + "messages": [{"role": "user", "content": f"[request {idx}] {PROMPT}"}], + "max_tokens": MAX_TOKENS, + "ignore_eos": True, # pin generated tokens == max_tokens (fixed denominator) + "temperature": 0.0, + "stream": False, + }).encode() + req = urllib.request.Request(url, data=body, headers={"Content-Type": "application/json"}) + t0 = time.time() + try: + with urllib.request.urlopen(req, timeout=300) as r: + d = json.loads(r.read()) + dt = time.time() - t0 + usage = d.get("usage") + if not usage or "completion_tokens" not in usage: + # a 200 without usage is a malformed success — surface it, don't count 0 tokens + return (False, dt, 0, "200-no-usage") + ct = usage["completion_tokens"] + if ct != MAX_TOKENS: + # ignore_eos is supposed to pin completion_tokens == MAX_TOKENS ("fixed denominator + # by construction"). CHECK it rather than assert it in prose: an endpoint that + # accepts the field but honours it loosely would otherwise book a short generation + # as a success with a variable numerator, quietly deflating tok/s. + return (False, dt, ct, f"short-gen-{ct}") + return (True, dt, ct, 200) + except urllib.error.HTTPError as e: + return (False, time.time() - t0, 0, e.code) + except Exception as e: + return (False, time.time() - t0, 0, str(e)[:40]) + +def pct(xs, p): + if not xs: + return 0.0 + xs = sorted(xs) + k = (len(xs) - 1) * p / 100.0 + f = int(k) + return xs[f] if f + 1 >= len(xs) else xs[f] + (xs[f + 1] - xs[f]) * (k - f) + +def sweep(conc, url, nreq_mult, base_idx): + # base_idx makes the prompt index MONOTONIC across levels: with a per-level range(n) + # the index restarts at 0 every level, so a prefix/KV cache warmed by level i serves + # level i+1's prefills for free — defeating the "unique prefix per request" control the + # methodology relies on. A run-global counter keeps uniqueness across the axis the sweep + # exists to compare, not just within a level. + n = conc * nreq_mult + t0 = time.time() + lat, toks, ok, codes = [], [], 0, {} + with ThreadPoolExecutor(max_workers=conc) as ex: + futs = [ex.submit(one_request, url, base_idx + i) for i in range(n)] + for fu in as_completed(futs): + success, dt, ct, code = fu.result() + if success: + ok += 1 + lat.append(dt); toks.append(ct) # successes only: failures must not + # deflate percentiles or ride in the throughput denominator + codes[str(code)] = codes.get(str(code), 0) + 1 + wall = time.time() - t0 + total_tok = sum(toks) + return { + "conc": conc, "n": n, "ok": ok, "wall_s": round(wall, 2), + "total_out_tok": total_tok, + "agg_tok_s": round(total_tok / wall, 1) if wall > 0 else 0, + "lat_p50_s": round(pct(lat, 50), 2), "lat_p90_s": round(pct(lat, 90), 2), + "lat_p99_s": round(pct(lat, 99), 2), "lat_max_s": round(max(lat), 2) if lat else 0, + "codes": codes, + } + +if __name__ == "__main__": + ap = argparse.ArgumentParser(description="trtllm-serve NcclEP live-serve throughput/latency probe") + ap.add_argument("--url", default=URL, help="chat-completions endpoint (default: 127.0.0.1 loopback)") + ap.add_argument("--out", default=None, help="write one JSON object per concurrency level to this path (JSONL)") + ap.add_argument("--requests-per-level-mult", type=int, default=NREQ_MULT, + help=f"requests per level = concurrency x this (default {NREQ_MULT})") + args = ap.parse_args() + # floor the multiplier: at 0 every level sends n=0, the final all(ok==n) gate is 0==0 + # (vacuously true), and the probe exits 0 having never touched the endpoint — which + # silently breaks the "exit code is the sole pass/fail authority" contract. + if args.requests_per_level_mult < 1: + ap.error("--requests-per-level-mult must be >= 1 (0 sends no requests and false-passes)") + + print("=== trtllm-serve NcclEP live-serve throughput probe ===") + print(f"url={args.url} model={MODEL} max_tokens={MAX_TOKENS} (ignore_eos) " + f"requests/level={args.requests_per_level_mult}x concurrency warmup={WARMUP}") + print(f"{'conc':>5} {'n':>5} {'ok':>5} {'wall_s':>7} {'out_tok':>8} {'agg_tok/s':>10} {'p50_s':>7} {'p90_s':>7} {'p99_s':>7} codes") + rows = [] + # run-global monotonic index (warmup requests consume the first WARMUP ids so the first + # measured request is still unique against them too). + base_idx = 0 + # discarded warmup so level 0 does not absorb lazy CUDA load / tokenizer init as ~20% of + # its sample — fired serially, results ignored, but the ids are still spent (uniqueness). + for _ in range(WARMUP): + one_request(args.url, base_idx) + base_idx += 1 + out_fh = open(args.out, "w") if args.out else None + try: + for c in CONCURRENCIES: + r = sweep(c, args.url, args.requests_per_level_mult, base_idx) + base_idx += r["n"] + rows.append(r) + print(f"{r['conc']:>5} {r['n']:>5} {r['ok']:>5} {r['wall_s']:>7} {r['total_out_tok']:>8} " + f"{r['agg_tok_s']:>10} {r['lat_p50_s']:>7} {r['lat_p90_s']:>7} {r['lat_p99_s']:>7} {r['codes']}") + if out_fh: + out_fh.write(json.dumps(r) + "\n"); out_fh.flush() + finally: + if out_fh: + out_fh.close() + print("JSON_ROWS=" + json.dumps(rows)) + # fail-loud contract: EVERY request at EVERY level must have succeeded — one dead + # backend at one level is a failed run, not a footnote (matches benchmark.sh's gate) + sys.exit(0 if all(rw["ok"] == rw["n"] for rw in rows) else 1) diff --git a/examples/inference/tensorrt-llm/nccl-ep-efa/recipe/probe_nccl_ep.py b/examples/inference/tensorrt-llm/nccl-ep-efa/recipe/probe_nccl_ep.py new file mode 100644 index 000000000..ef95b2f29 --- /dev/null +++ b/examples/inference/tensorrt-llm/nccl-ep-efa/recipe/probe_nccl_ep.py @@ -0,0 +1,298 @@ +#!/usr/bin/env python3 +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. SPDX-License-Identifier: MIT-0 +"""probe_nccl_ep.py — does a real serve's selection path hand back an NcclEP, and +does that instance move bytes cross-node on EFA? + +A serve never constructs NcclEP itself — it calls +CommunicationFactory.create_strategy(model_config, ...), which reads +TRTLLM_FORCE_COMM_METHOD, runs TRT-LLM's own _get_nccl_ep_unavailable_reason +feasibility gate, and derives every NcclEP ctor arg from model_config. So this +probe drives THAT entrypoint, with a ModelConfig stand-in carrying the same +fields a serve would (mapping with enable_attention_dp + moe_tp_size=1, +pretrained_config.hidden_size, torch_dtype, max_num_tokens, ...), then runs +dispatch/combine through whatever the factory returned and checks the combined +output against a CPU-derived oracle. + +Runs under torchrun (see run-kernel-test.sh) so the strategy the factory selects +is exercised ACROSS the node boundary, not just instantiated — the defect class +this exists to catch is a first-dispatch illegal memory access that +instantiation alone never reaches. +""" +import os +import sys +import traceback + +import torch +import torch.distributed as dist + +# RANK/WORLD/LOCAL are the one required (torchrun-set) trio; they are read inside +# main() — under the bottom try/except — so a run outside torchrun surfaces as PROBE-EXC +# rather than a bare KeyError traceback. -1 sentinels keep log()/the except usable if the +# read never happened. The EP_* knobs below use .get() defaults, so they cannot KeyError. +RANK = WORLD = LOCAL = -1 + +E = int(os.environ.get("EP_EXPERTS", "128")) +TOK = int(os.environ.get("EP_TOKENS", "128")) +H = int(os.environ.get("EP_HIDDEN", "2048")) +K = int(os.environ.get("EP_TOPK", "8")) +NUM_QP = int(os.environ.get("EP_NUM_QP", "32")) + + +def log(msg): + print(f"[rank{RANK}] {msg}", flush=True) + + +def install_mpi_shim(): + """mpi4py-free stand-in for mpi_comm(), used only for NCCL unique-id bootstrap. + + trtllm-serve runs under mpirun and gets a real MPI communicator; under + torchrun there is none, so the handful of collective ops TRT-LLM's NcclEP + bootstrap needs (bcast/allgather/barrier of the NCCL unique id) are backed + by torch.distributed instead. This shims the transport of the bootstrap + only — the EP data path under test is untouched. + """ + import tensorrt_llm._utils as tu + + class _Shim: + def Get_rank(self): + return dist.get_rank() + + def Get_size(self): + return dist.get_world_size() + + def Split(self, color, key): + # Returns the world communicator, valid ONLY while ep_size == world_size (this + # probe's shape — see build_model_config's moe_ep_size=WORLD): the EP group then + # IS the world group. At any shape where ep_size < world_size this would silently + # build the wrong group; the probe would need a real color-keyed split first. + return self + + def bcast(self, obj, root=0): + box = [obj if dist.get_rank() == root else None] + dist.broadcast_object_list(box, src=root) + return box[0] + + def allgather(self, obj): + out = [None] * dist.get_world_size() + dist.all_gather_object(out, obj) + return out + + def barrier(self): + dist.barrier() + + def Free(self): + pass + + shim = _Shim() + tu.mpi_comm = lambda: shim + tu.mpi_rank = lambda: dist.get_rank() + tu.mpi_world_size = lambda: dist.get_world_size() + tu.mpi_barrier = lambda: dist.barrier() + import tensorrt_llm + + for mod in (tensorrt_llm, tu): + for name in ("mpi_comm", "mpi_rank", "mpi_world_size"): + if hasattr(mod, name): + setattr(mod, name, getattr(tu, name)) + + +def build_model_config(mapping): + """A ModelConfig with exactly the attributes the factory reads. + + Using a stand-in rather than a real ModelConfig keeps the probe independent + of a checkpoint download, while the attribute set is taken from the factory + source (mapping, pretrained_config.hidden_size, torch_dtype, quant_config, + max_num_tokens, moe_max_num_tokens, use_cuda_graph, + use_low_precision_moe_combine, moe_load_balancer). + """ + + class _Pretrained: + hidden_size = H + + class _MC: + pass + + mc = _MC() + mc.mapping = mapping + mc.pretrained_config = _Pretrained() + mc.torch_dtype = torch.bfloat16 + mc.quant_config = None + mc.max_num_tokens = TOK + mc.moe_max_num_tokens = None + mc.use_cuda_graph = False + mc.use_low_precision_moe_combine = False + mc.moe_load_balancer = None + return mc + + +def main(): + global RANK, WORLD, LOCAL + # Read the torchrun-set trio here (not at module scope) so a run outside torchrun + # surfaces as PROBE-EXC via the bottom handler, not a bare KeyError at import. + RANK = int(os.environ["RANK"]) + WORLD = int(os.environ["WORLD_SIZE"]) + LOCAL = int(os.environ["LOCAL_RANK"]) + torch.cuda.set_device(LOCAL) + dist.init_process_group("nccl", rank=RANK, world_size=WORLD) + dev = torch.device("cuda", LOCAL) + os.environ.setdefault("NCCL_EP_NUM_QP_PER_RANK", str(NUM_QP)) + log(f"probe start world={WORLD} algo={os.environ.get('TRTLLM_NCCL_EP_ALGO', '(upstream default)')}") + + # EP requires the experts to shard evenly across ranks. At an indivisible world size + # (e.g. a 3-node run, WORLD=24, E=128) the E//WORLD floor orphans the tail experts: + # tokens still route to them, their contributions never enter moe_out, and the probe + # would report a false PROBE-MISMATCH on a healthy fabric. E and WORLD are rank-invariant + # (env default + torchrun), so this fires identically on every rank — no split-hang — and + # also guards the expert_size_per_partition=E//WORLD arg handed to the factory below. + assert E % WORLD == 0, ( + f"EP_EXPERTS={E} must be divisible by world size {WORLD} " + f"({E % WORLD} experts would be orphaned) — pick nodes×8 that divides {E}" + ) + + install_mpi_shim() + + from tensorrt_llm._torch.modules.fused_moe.communication.communication_factory import ( + CommunicationFactory, + ) + from tensorrt_llm._torch.modules.fused_moe.communication.nccl_ep import NcclEP + from tensorrt_llm.mapping import Mapping + + # enable_attention_dp + moe_tp_size=1 are what a wide-EP serve runs; without + # them create_strategy short-circuits to None / AllGatherReduceScatter BEFORE + # it ever reads TRTLLM_FORCE_COMM_METHOD (same gate serve.sh satisfies via + # its extra_llm_api_options YAML). + mapping = Mapping( + world_size=WORLD, + rank=RANK, + gpus_per_node=8, + tp_size=WORLD, + moe_tp_size=1, + moe_ep_size=WORLD, + enable_attention_dp=True, + ) + mc = build_model_config(mapping) + + reason = CommunicationFactory._get_nccl_ep_unavailable_reason( + torch.bfloat16, None, E, H, TOK, None, K + ) + log(f"PROBE-GATE unavailable_reason={reason!r}") + + # The real serve entrypoint. Reads TRTLLM_FORCE_COMM_METHOD itself. + strategy = CommunicationFactory.create_strategy( + model_config=mc, + num_experts=E, + num_slots=E, + top_k=K, + expert_size_per_partition=E // WORLD, + ) + log(f"PROBE-SELECTED type={type(strategy).__name__}") + # Agree the selection verdict across ranks BEFORE any rank returns: the factory's + # feasibility gate (_get_nccl_ep_unavailable_reason's per-device shared-memory check) + # and NcclEP.__init__'s is_nccl_ep_installed() are per-rank evaluations that a + # heterogeneous node or a divergent library path can split. A bare per-rank return + # would leave the survivors blocked in the bootstrap collectives until torchrun's + # timeout — same all-reduce(MIN) contract the combine verdict below uses. + sel_ok = torch.tensor([1 if isinstance(strategy, NcclEP) else 0], device=dev, dtype=torch.int32) + dist.all_reduce(sel_ok, op=dist.ReduceOp.MIN) + if int(sel_ok.item()) != 1: + log(f"PROBE-FAIL factory returned {type(strategy).__name__}, not NcclEP (on this or a peer rank)") + return 5 + + ctx = strategy._get_context() + # NcclEpContext on 1.3.0rc24 exposes `layout` but NOT `_ep_algorithm` — the algorithm is + # not stored on the context in the unpatched factory (upstream hardcodes LOW_LATENCY). A + # getattr-guard keeps the diagnostic from AttributeError-ing on a GIN-working substrate: + # when the attr is absent, report the requested algo (env), which == what runs on the + # unpatched default path. Prefer the real attr if a future base exposes it. + _algo = getattr(ctx, "_ep_algorithm", None) + algo_name = _algo.name if _algo is not None else os.environ.get("TRTLLM_NCCL_EP_ALGO", "LOW_LATENCY").upper() + log(f"PROBE-CONFIG algorithm={algo_name} layout={ctx.layout.name}") + + # Upstream default is LOW_LATENCY+RANK_MAJOR; the algorithm knob is only + # honored on an image built with APPLY_HT_FLAT_PATCH=1 (TensorRT-LLM PR + # #17715). If HT was explicitly requested, silently probing LL instead + # would be a false pass — fail loud. + want = os.environ.get("TRTLLM_NCCL_EP_ALGO", "LOW_LATENCY").upper() + if want in ("HIGH_THROUGHPUT", "HT"): + # rc24 does not expose the algorithm as a context attr; layout IS observable + # (RANK_MAJOR default, FLAT once PR #17715's HT/FLAT patch is baked). Prefer the + # algo attr if a future base exposes it, else fall back to the layout==FLAT proof. + algo_ok = (_algo.name == "HIGH_THROUGHPUT") if _algo is not None else (ctx.layout.name == "FLAT") + if not algo_ok: + log("PROBE-FAIL env requested HIGH_THROUGHPUT but the factory selected " + f"{algo_name}+{ctx.layout.name} — unpatched image? " + "(build with APPLY_HT_FLAT_PATCH=1)") + return 6 + + # Exercise the factory-selected instance cross-node — selection alone does not + # prove the group works, and the whole defect class here is a first-dispatch IMA. + g = torch.Generator(device="cpu").manual_seed(4242 + RANK) + # Snapshot the oracle inputs on CPU FIRST, then hand independent device copies to + # dispatch. expect (below) is derived from these CPU snapshots, so a kernel that + # mutated its caller-owned inputs in place could not corrupt the reference and the + # result identically and still pass — the docstring's "CPU-derived oracle" is now literal. + x_cpu = (torch.randn(TOK, H, generator=g, dtype=torch.float32) / 8.0).to(torch.bfloat16) + slots_cpu = torch.stack([torch.randperm(E, generator=g)[:K] for _ in range(TOK)]).to(torch.int32) + scales_cpu = torch.rand(TOK, K, generator=g, dtype=torch.float32) + x = x_cpu.to(dev) + slots = slots_cpu.to(dev) + scales = scales_cpu.to(dev) + + recv_hs, _, recv_slots, recv_scales = strategy.dispatch( + hidden_states=x, + hidden_states_sf=None, + token_selected_slots=slots, + token_final_scales=scales, + all_rank_num_tokens=[TOK] * WORLD, + ) + torch.cuda.synchronize() + log(f"PROBE-DISPATCH recv_hs={tuple(recv_hs.shape)}") + + # Local "expert compute" stand-in with a closed-form oracle: expert e scales its + # tokens by (e+1)/E, so the combined output must equal + # x * sum_k(((slot_k+1)/E) * scale_k) — checkable per token without real weights. + num_local = E // WORLD + lo, hi = RANK * num_local, RANK * num_local + num_local + moe_out = torch.zeros_like(recv_hs, dtype=torch.bfloat16) + sl = recv_slots.to(torch.int64) + for kk in range(recv_slots.shape[1]): + e = sl[:, kk] + m = (e >= lo) & (e < hi) + if not bool(m.any()): + continue + f = ((e[m].to(torch.float32) + 1.0) / E).unsqueeze(1) + w = recv_scales[m, kk].unsqueeze(1).to(torch.float32) + moe_out[m] += (recv_hs[m].to(torch.float32) * f * w).to(torch.bfloat16) + + combined = strategy.combine(moe_out, all_rank_max_num_tokens=TOK) + torch.cuda.synchronize() + + # Oracle from the CPU snapshots (not the tensors dispatch received), moved to device + # only for the comparison against the device-resident combined output. + fac = ((slots_cpu.to(torch.float32) + 1.0) / E) * scales_cpu + expect = (x_cpu.to(torch.float32) * fac.sum(1, keepdim=True)).to(dev) + got = combined.to(torch.float32) + relmax = ((got - expect).abs().max() / expect.abs().max().clamp_min(1e-6)).item() + nz = int((got.abs().sum(1) > 0).sum()) + ok = (nz == TOK) and (relmax < 0.05) + log(f"PROBE-COMBINE nonzero={nz}/{TOK} relmax={relmax:.4g} {'OK' if ok else 'MISMATCH'}") + + # Every rank must agree — one silently-corrupted rank is a failed run. + flag = torch.tensor([1 if ok else 0], device=dev, dtype=torch.int32) + dist.all_reduce(flag, op=dist.ReduceOp.MIN) + verdict = "PROBE-PASS" if int(flag.item()) == 1 else "PROBE-MISMATCH" + log(f"{verdict} factory-selected NcclEP {algo_name}+{ctx.layout.name} world={WORLD}") + + strategy.destroy() + dist.barrier() + log("PROBE-DONE") + return 0 if int(flag.item()) == 1 else 4 + + +if __name__ == "__main__": + try: + sys.exit(main()) + except Exception: + traceback.print_exc() + print(f"[rank{RANK}] PROBE-EXC", flush=True) + sys.exit(1) diff --git a/examples/inference/tensorrt-llm/nccl-ep-efa/recipe/run-kernel-test.sh b/examples/inference/tensorrt-llm/nccl-ep-efa/recipe/run-kernel-test.sh new file mode 100755 index 000000000..ef944e324 --- /dev/null +++ b/examples/inference/tensorrt-llm/nccl-ep-efa/recipe/run-kernel-test.sh @@ -0,0 +1,87 @@ +#!/bin/bash +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. SPDX-License-Identifier: MIT-0 +# run-kernel-test.sh — prove the factory-selected NcclEP actually moves bytes over EFA +# BEFORE any model load. This is the "smoke-test the transport" gate: it drives TRT-LLM's +# own CommunicationFactory.create_strategy() + dispatch/combine across the nodes with the +# SAME NcclEP/EFA env the serve uses (probe_nccl_ep.py), and exits non-zero on failure. +# Cheap, and — unlike a node-spanning serve, which needs mpirun (see the kubernetes +# manifest header) — this torchrun probe DOES cross the node boundary. +# +# leader: run-kernel-test.sh leader +# worker: run-kernel-test.sh worker # node-rank 1,2,3 for the worker nodes +# +# Env (shared with serve.sh): NNODES (default 2), GPUS_PER_NODE (default 8), +# TRTLLM_NCCL_EP_ALGO (default LOW_LATENCY — see the patched-image guard below). +set -uo pipefail + +# NOTE: keep the usage message brace-free. A literal '}' inside a ${1:?...} message (e.g. +# "{leader|worker}") closes the parameter expansion early, so ROLE captures the junk tail and +# the case below never matches -> FATAL on every invocation. Check role in a separate statement. +ROLE="${1:-}"; : "${ROLE:?usage: run-kernel-test.sh leader-or-worker [node-rank]}" +case "$ROLE" in leader|worker) ;; *) echo "FATAL: unrecognized role '$ROLE' (expected leader|worker)"; exit 2 ;; esac +LEADER_IP="${2:?need leader ip}" +if [ "$ROLE" = "worker" ]; then NODE_RANK_ARG="${3:?worker requires an explicit node-rank (1,2,...)}"; else NODE_RANK_ARG=0; fi +NNODES="${NNODES:-2}" +GPUS_PER_NODE="${GPUS_PER_NODE:-8}" +PROBE="/opt/probe_nccl_ep.py" +[ -f "$PROBE" ] || { echo "FAIL: $PROBE not in image — rebuild from the Dockerfile"; exit 3; } + +# ---- NCCL-GIN proxy + EFA contract, VERBATIM from serve.sh (same transport under test) ---- +export NCCL_GIN_TYPE=2 NCCL_GIN_ENABLE=1 # 2 = CPU-proxy GIN, the EFA-viable GIN mode +export NCCL_CUMEM_ENABLE=1 NCCL_NVLS_ENABLE=0 NCCL_IGNORE_DISABLED_P2P=1 +export FI_PROVIDER=efa FI_EFA_USE_DEVICE_RDMA=1 FI_EFA_ENABLE_SHM_TRANSFER=0 FI_EFA_FORK_SAFE=1 +export OFI_NCCL_PROTOCOL=RDMA +export NCCL_NET_PLUGIN=/opt/aws-ofi-nccl/lib/libnccl-net-ofi.so +# ---- NcclEP selection (same knobs the serve exports) ---- +export TRTLLM_FORCE_COMM_METHOD=NCCL_EP +export NCCL_EP_NUM_QP_PER_RANK="${NCCL_EP_NUM_QP_PER_RANK:-32}" +export ENABLE_CONFIGURABLE_MOE=1 # match serve.sh: the correctness gate must exercise the + # SAME configurable-MoE path the serve runs, not a variant +export TLLM_LOG_LEVEL=info # TLLM_, not TRTLLM_ — the latter is silently ignored (README trap 1) +export TRTLLM_NCCL_EP_ALGO="${TRTLLM_NCCL_EP_ALGO:-LOW_LATENCY}" +if [ "$TRTLLM_NCCL_EP_ALGO" != "LOW_LATENCY" ] && [ ! -f /opt/.ht-flat-patch-applied ]; then + echo "FATAL: TRTLLM_NCCL_EP_ALGO=$TRTLLM_NCCL_EP_ALGO needs an image built with" + echo "APPLY_HT_FLAT_PATCH=1 (TensorRT-LLM PR#17715). Unpatched upstream ignores the knob" + echo "and runs LOW_LATENCY+RANK_MAJOR — refusing rather than probing the wrong algorithm." + exit 4 +fi +export NCCL_DEBUG=${KERNEL_TEST_NCCL_DEBUG:-INFO} # INFO so the efa-direct banner prints = transport proof + +# ---- lib path: pip NCCL 2.30.4 first (the container's baked NCCL must NOT win) ---- +# Fail loud rather than swallow (see serve.sh): an empty NCCL_LIB would silently drop the +# "pip NCCL wins" contract (README trap 3) and prepend an empty LD_LIBRARY_PATH element. +# The :? guard fires even though set -e is off here (parameter expansion exits regardless). +NCCL_LIB="$(python3 -c 'import importlib.util,os +s=importlib.util.find_spec("nvidia.nccl") +p=(s.submodule_search_locations[0] if s and s.submodule_search_locations else None) +print(os.path.join(p,"lib") if p else "")')" +: "${NCCL_LIB:?could not locate the pip nvidia-nccl-cu13 lib dir — NcclEP needs it first on LD_LIBRARY_PATH (README trap 3)}" +export LD_LIBRARY_PATH="${NCCL_LIB}:/opt/aws-ofi-nccl/lib:/opt/amazon/efa/lib:/usr/local/cuda/lib64:${LD_LIBRARY_PATH:-}" +export PATH=/opt/amazon/efa/bin:$PATH + +NODE_RANK="$NODE_RANK_ARG"; [ "$ROLE" = "leader" ] && NODE_RANK=0 +echo "===== NcclEP kernel smoke: role=$ROLE node_rank=$NODE_RANK nnodes=$NNODES gpus/node=$GPUS_PER_NODE algo=$TRTLLM_NCCL_EP_ALGO leader=$LEADER_IP $(hostname) $(date -u +%FT%TZ) =====" + +set -o pipefail +# 8 torchrun procs per node (one per GPU): probe_nccl_ep.py is a per-rank script that +# reads RANK/WORLD_SIZE/LOCAL_RANK — 2 nodes x 8 = the 16-rank EP16 shape the measured +# runs used. The probe's MPI shim backs TRT-LLM's bootstrap collectives with +# torch.distributed, which is why this crosses nodes where mpirun cannot. +torchrun --nnodes="$NNODES" --nproc-per-node="$GPUS_PER_NODE" --node-rank="$NODE_RANK" \ + --master-addr="$LEADER_IP" --master-port=29501 "$PROBE" 2>&1 | tee /tmp/kernel-test.$NODE_RANK.log +rc=${PIPESTATUS[0]} + +# fail-loud contract: gate on the exit code (the probe all-reduces a MIN verdict — one bad +# rank fails every rank), then separately require the EFA transport banner. +if [ "$rc" -eq 0 ]; then + # confirm EFA (not TCP/SHM) actually carried it — only provider-specific banners count; + # a bare "NET/OFI" line also prints for tcp;ofi_rxm fallback, the exact case to rule out + if grep -qiE "efa-direct|Selected Provider is efa" /tmp/kernel-test.$NODE_RANK.log; then + echo "KERNEL-TEST PASS (node_rank=$NODE_RANK) — factory-selected NcclEP dispatch/combine over EFA verified" + exit 0 + fi + echo "KERNEL-TEST INCONCLUSIVE: probe passed but no EFA-provider banner in log — confirm transport before trusting" + exit 2 +fi +echo "KERNEL-TEST FAIL (node_rank=$NODE_RANK, torchrun rc=$rc) — see /tmp/kernel-test.$NODE_RANK.log" +exit 1 diff --git a/examples/inference/tensorrt-llm/nccl-ep-efa/recipe/serve.sh b/examples/inference/tensorrt-llm/nccl-ep-efa/recipe/serve.sh new file mode 100755 index 000000000..5111adb39 --- /dev/null +++ b/examples/inference/tensorrt-llm/nccl-ep-efa/recipe/serve.sh @@ -0,0 +1,128 @@ +#!/bin/bash +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. SPDX-License-Identifier: MIT-0 +# serve.sh — trtllm-serve with the NcclEP MoE all-to-all backend over AWS EFA +# (NCCL-GIN CPU-proxy). SINGLE-NODE (tp_size == ep_size == GPUs on the node). +# +# Why single-node: a node-spanning trtllm-serve needs mpirun, and mpirun's OOB cannot +# route between EKS pods — AWS VPC-CNI gives each pod a /32 eth0, so OMPI's +# opal_net_samenetwork() never matches a peer address and the HNP reports "no route to" +# the remote orted. torchrun does not share that limitation, which is why the CROSS-NODE +# proof lives in run-kernel-test.sh (16-rank factory/dispatch/combine probe), while the +# served-completion proof is this script at EP8. See README "Known limitations". +# +# serve.sh # in-pod (the kubernetes manifest runs this on ordinal 0) +# +# Knobs (env; defaults = the measured EP8 shape): +# SERVE_MODEL=... any MoE whose routed-expert count % SERVE_EP == 0 (preflighted below) +# SERVE_TP=8 SERVE_EP=8 single-node: both == GPUs on the node +# TRTLLM_NCCL_EP_ALGO=LOW_LATENCY non-default values need the patched image (guard below) +set -euo pipefail # -e: preflight failures below must STOP the launch, not fall through to trtllm-serve + +# ---- NCCL-GIN proxy + EFA env contract (identical to the measured runs + deploy YAML) ---- +export NCCL_GIN_TYPE=2 NCCL_GIN_ENABLE=1 # 2 = CPU-proxy GIN, the EFA-viable GIN mode +export NCCL_CUMEM_ENABLE=1 NCCL_NVLS_ENABLE=0 NCCL_IGNORE_DISABLED_P2P=1 +export FI_PROVIDER=efa FI_EFA_USE_DEVICE_RDMA=1 FI_EFA_ENABLE_SHM_TRANSFER=0 FI_EFA_FORK_SAFE=1 +export OFI_NCCL_PROTOCOL=RDMA +export NCCL_NET_PLUGIN=/opt/aws-ofi-nccl/lib/libnccl-net-ofi.so +export NCCL_DEBUG=${SERVE_NCCL_DEBUG:-WARN} # INFO to see the efa-direct banner proof + +# ---- NcclEP selection — the point of this sample ---- +# TRTLLM_FORCE_COMM_METHOD=NCCL_EP makes CommunicationFactory.create_strategy() return +# NcclEP — but ONLY once attention DP is on (the extra_llm_api_options YAML below); +# without it the factory short-circuits to None BEFORE reading this env. +export TRTLLM_FORCE_COMM_METHOD=NCCL_EP +export NCCL_EP_NUM_QP_PER_RANK="${NCCL_EP_NUM_QP_PER_RANK:-32}" # QPs per rank — part of benchmark provenance +export ENABLE_CONFIGURABLE_MOE=1 +# TLLM_LOG_LEVEL (not TRTLLM_LOG_LEVEL — that one is silently ignored) is what +# tensorrt_llm/logger.py reads; without info the "NCCL EP group created ..." line that +# constitutes engagement PROOF is suppressed at the default 'error'. README trap 1. +export TLLM_LOG_LEVEL=info + +# ---- algorithm knob guard: honored ONLY on a patched image (TensorRT-LLM PR#17715) ---- +# Unpatched upstream hardcodes LOW_LATENCY+RANK_MAJOR (the measured-clean baseline on this +# substrate). Refusing beats silently serving a different algorithm than the one requested. +export TRTLLM_NCCL_EP_ALGO="${TRTLLM_NCCL_EP_ALGO:-LOW_LATENCY}" +if [ "$TRTLLM_NCCL_EP_ALGO" != "LOW_LATENCY" ] && [ ! -f /opt/.ht-flat-patch-applied ]; then + echo "REFUSING TRTLLM_NCCL_EP_ALGO=$TRTLLM_NCCL_EP_ALGO on an unpatched image: upstream" + echo "ignores the knob and runs LOW_LATENCY+RANK_MAJOR regardless. Build with" + echo "APPLY_HT_FLAT_PATCH=1 (bakes TensorRT-LLM PR#17715) to make it selectable." + exit 4 +fi + +# ---- HF cache on the pod-local writable volume (NOT the image layer) ---- +export HF_HOME=${HF_HOME:-/work/hf} +export HUGGINGFACE_HUB_CACHE=${HUGGINGFACE_HUB_CACHE:-$HF_HOME/hub} +mkdir -p "$HUGGINGFACE_HUB_CACHE" + +# ---- lib path: pip NCCL 2.30.4 first (the container's baked NCCL must NOT win) ---- +# No `2>/dev/null || true`: that swallowed every failure, left NCCL_LIB empty, silently +# dropped the "pip NCCL wins" contract (README trap 3), and put a leading empty element on +# LD_LIBRARY_PATH (the loader then searches $PWD first). Fail loud instead — the whole +# sample depends on this resolving, and the downstream symptom otherwise misdirects +# ("nccl-ep is not installed" points at nccl4py when the real cause is a lost lib path). +NCCL_LIB="$(python3 -c 'import importlib.util,os +s=importlib.util.find_spec("nvidia.nccl") +p=(s.submodule_search_locations[0] if s and s.submodule_search_locations else None) +print(os.path.join(p,"lib") if p else "")')" +: "${NCCL_LIB:?could not locate the pip nvidia-nccl-cu13 lib dir — NcclEP needs it first on LD_LIBRARY_PATH (README trap 3)}" +export LD_LIBRARY_PATH="${NCCL_LIB}:/opt/aws-ofi-nccl/lib:/opt/amazon/efa/lib:/usr/local/cuda/lib64:${LD_LIBRARY_PATH:-}" +export PATH=/opt/amazon/efa/bin:$PATH + +# ---- scale + model (defaults = the measured single-node EP8 shape) ---- +SERVE_MODEL="${SERVE_MODEL:-Qwen/Qwen3-30B-A3B}" +# --trust_remote_code executes arbitrary repo Python in a privileged root container, so pin +# WHAT executes: SERVE_REVISION (a commit SHA/tag) makes the model an immutable artifact like +# every other pin here, instead of "whatever main points at today". Empty = HF default branch. +SERVE_REVISION="${SERVE_REVISION:-}" +SERVE_TP="${SERVE_TP:-8}" +SERVE_EP="${SERVE_EP:-8}" +SERVE_HOST="${SERVE_HOST:-0.0.0.0}" +SERVE_PORT="${SERVE_PORT:-8000}" +SERVE_MAX_BATCH_SIZE="${SERVE_MAX_BATCH_SIZE:-4}" +SERVE_MAX_NUM_TOKENS="${SERVE_MAX_NUM_TOKENS:-2048}" +SERVE_MAX_SEQ_LEN="${SERVE_MAX_SEQ_LEN:-2048}" +SERVE_KV_FRACTION="${SERVE_KV_FRACTION:-0.5}" + +# ---- EP-divisibility preflight (the one per-model gate): routed experts % EP == 0 ---- +# Qwen3 family=128 (ok 8/16/32), DeepSeek-V3/R1=256 (ok), Qwen1.5-MoE=60 (FAILS 8/16). +# Reads the model's config.json via huggingface_hub. +if [ "${SKIP_EP_PREFLIGHT:-0}" != "1" ]; then + python3 - "$SERVE_MODEL" "$SERVE_EP" "$SERVE_REVISION" <<'PY' || exit 3 +import json, sys +from huggingface_hub import hf_hub_download +model, ep = sys.argv[1], int(sys.argv[2]) +rev = sys.argv[3] or None # validate the SAME pinned revision the serve will run +cfg = json.load(open(hf_hub_download(model, "config.json", revision=rev))) +n = cfg.get("n_routed_experts") or cfg.get("num_experts") +assert n, f"{model}: no n_routed_experts/num_experts in config.json — not a routed-MoE model?" +assert n % ep == 0, (f"EP-DIVISIBILITY GATE FAILED: {model} has {n} routed experts, " + f"not divisible by EP={ep}. Pick an EP that divides {n} or another model.") +print(f"EP preflight OK: {model} routed_experts={n} % EP={ep} == 0 ({n//ep} experts/rank)") +PY +fi + +# ---- attention DP: REQUIRED for NcclEP (see the selection comment above) ---- +CFG=/tmp/extra-llm-api-config.yml +cat > "$CFG" <<'YML' +enable_attention_dp: true +YML + +echo "===== trtllm-serve model=$SERVE_MODEL tp=${SERVE_TP} ep=${SERVE_EP} algo=${TRTLLM_NCCL_EP_ALGO} qp/rank=${NCCL_EP_NUM_QP_PER_RANK} $(hostname) $(date -u +%FT%TZ) =====" +python3 -c "import tensorrt_llm; print('tensorrt_llm', tensorrt_llm.__version__)" +python3 -c "import nccl.ep; print('nccl.ep OK (nccl4py)')" + +# Engagement PROOF once up (requires TLLM_LOG_LEVEL=info above): +# grep "NCCL EP group created" +# -> "... layout=RANK_MAJOR, algorithm=LOW_LATENCY" on every rank (or FLAT/HIGH_THROUGHPUT +# on a patched image with TRTLLM_NCCL_EP_ALGO=HIGH_THROUGHPUT) +# --trust_remote_code: MoE models whose HF repos ship modeling code (DeepSeek family) need +# it; a SERVE_MODEL you do not trust should not be pointed at this serve. +exec trtllm-serve "$SERVE_MODEL" \ + ${SERVE_REVISION:+--revision "$SERVE_REVISION"} \ + --host "$SERVE_HOST" --port "$SERVE_PORT" \ + --tp_size "$SERVE_TP" --ep_size "$SERVE_EP" --pp_size 1 \ + --max_batch_size "$SERVE_MAX_BATCH_SIZE" --max_num_tokens "$SERVE_MAX_NUM_TOKENS" \ + --max_seq_len "$SERVE_MAX_SEQ_LEN" \ + --kv_cache_free_gpu_memory_fraction "$SERVE_KV_FRACTION" \ + --trust_remote_code \ + --extra_llm_api_options "$CFG" diff --git a/examples/inference/tensorrt-llm/nccl-ep-efa/recipe/verify-image.sh b/examples/inference/tensorrt-llm/nccl-ep-efa/recipe/verify-image.sh new file mode 100755 index 000000000..6a124a8b0 --- /dev/null +++ b/examples/inference/tensorrt-llm/nccl-ep-efa/recipe/verify-image.sh @@ -0,0 +1,53 @@ +#!/usr/bin/env bash +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. SPDX-License-Identifier: MIT-0 +# Smoke the EFA + NcclEP substrate in the built image BEFORE loading a model. Fails loud. +# Static image check: asserts what the image stages (libs, symbols, imports, patch-marker +# consistency). The live cross-node transport proof is run-kernel-test.sh, not this. +set -euo pipefail +IMG="${1:?usage: verify-image.sh }" + +# EFA device mapping: fi_info -p efa only resolves with the device visible in the +# container. On an EFA host, pass /dev/infiniband through; elsewhere fall back to a +# provider-compiled-in check (fi_info -l) and say so. +DEV_ARGS=() +HAVE_EFA_DEV=0 +if [ -d /dev/infiniband ]; then + DEV_ARGS=(-v /dev/infiniband:/dev/infiniband --device=/dev/infiniband) + HAVE_EFA_DEV=1 +fi + +docker run --rm --gpus all "${DEV_ARGS[@]}" -e HAVE_EFA_DEV="${HAVE_EFA_DEV}" "${IMG}" bash -lc ' + set -euo pipefail + if [ "${HAVE_EFA_DEV}" = "1" ]; then + echo "== fi_info efa (live fabric) ==" + [ "$(/opt/amazon/efa/bin/fi_info -p efa | grep -c "fabric: efa-direct")" -ge 1 ] || { echo "FAIL: no efa-direct"; exit 1; } + else + echo "== fi_info efa (no /dev/infiniband on this host — checking provider is compiled in only) ==" + /opt/amazon/efa/bin/fi_info -l | grep -qi "efa" || { echo "FAIL: efa provider not in libfabric"; exit 1; } + echo " (run this script on an EFA host for the live efa-direct fabric check)" + fi + echo "== single libnccl 2.30.4 wins (path + version string — the NGC base bakes an older NCCL in a DIFFERENT dir, so path matters) ==" + # draining form (awk NR==1, no `head`): `head -1` closes the pipe after one line, so grep + # takes SIGPIPE-141 under pipefail — worst exactly when there are 2+ libnccl entries (the + # baked-shadows-pip case this check exists to catch). Same trap the Dockerfile calls out. + NCCL_SO=$(ldconfig -p | grep "libnccl.so.2 " | awk "NR==1{print \$NF}") + echo "$NCCL_SO" | grep -q "nvidia/nccl" || { echo "FAIL: baked libnccl shadows pip ($NCCL_SO)"; exit 1; } + [ "$(strings "$NCCL_SO" | grep -c "NCCL version 2.30.4")" -ge 1 ] || { echo "FAIL: $NCCL_SO is not 2.30.4 — NcclEP is_nccl_ep_installed() gates on >= 2.30.4"; exit 1; } + echo "== GIN plugin symbol =="; [ "$(nm -D /opt/aws-ofi-nccl/lib/libnccl-net-ofi.so | grep -c ncclGinPlugin)" -ge 1 ] || { echo "FAIL: no ncclGinPlugin"; exit 1; } + echo "== nccl.ep python package (nccl4py) ==" + python3 -c "import nccl.ep; print(\"nccl.ep at\", nccl.ep.__file__)" || { echo "FAIL: import nccl.ep"; exit 1; } + echo "== tensorrt_llm + the NcclEP factory module ==" + python3 -c "import tensorrt_llm; print(\"tensorrt_llm\", tensorrt_llm.__version__)" + python3 -c "from tensorrt_llm._torch.modules.fused_moe.communication.communication_factory import CommunicationFactory; from tensorrt_llm._torch.modules.fused_moe.communication.nccl_ep import NcclEP; print(\"factory + NcclEP import OK\")" \ + || { echo "FAIL: NcclEP factory modules missing — wrong base tag? (GA v1.2.1 has no NcclEP backend)"; exit 1; } + echo "== patch-marker consistency (marker and site-packages must agree) ==" + if [ -f /opt/.ht-flat-patch-applied ]; then + SP=$(python3 -c "import tensorrt_llm, pathlib; print(pathlib.Path(tensorrt_llm.__file__).parent)") + [ "$(grep -rc TRTLLM_NCCL_EP_ALGO "$SP/_torch/modules/fused_moe/communication/" 2>/dev/null | awk -F: "{s+=\$2} END {print s+0}")" -ge 1 ] \ + || { echo "FAIL: marker says patched but TRTLLM_NCCL_EP_ALGO gate not in site-packages"; exit 1; } + echo " patched image (PR#17715 baked — algorithm/layout selectable)" + else + echo " unpatched baseline (upstream LOW_LATENCY+RANK_MAJOR — the measured-clean default)" + fi + echo "ALL CHECKS PASS" +' diff --git a/examples/inference/tensorrt-llm/nccl-ep-efa/setup/build-push.sh b/examples/inference/tensorrt-llm/nccl-ep-efa/setup/build-push.sh new file mode 100755 index 000000000..ff0b0fb68 --- /dev/null +++ b/examples/inference/tensorrt-llm/nccl-ep-efa/setup/build-push.sh @@ -0,0 +1,14 @@ +#!/usr/bin/env bash +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. SPDX-License-Identifier: MIT-0 +set -euo pipefail +cd "$(dirname "$0")/.." +[ -f setup/env_vars ] || { echo "FATAL: cp setup/env_vars.example setup/env_vars and edit it first"; exit 2; } +source setup/env_vars +: "${REGISTRY:?set REGISTRY in setup/env_vars}" +IMG="${REGISTRY}/${IMAGE_NAME}:${IMAGE_TAG}" +aws ecr get-login-password --region "${AWS_REGION}" | docker login --username AWS --password-stdin "${REGISTRY}" +DOCKER_BUILDKIT=1 docker build \ + ${APPLY_HT_FLAT_PATCH:+--build-arg APPLY_HT_FLAT_PATCH="${APPLY_HT_FLAT_PATCH}"} \ + -t "${IMG}" . +docker push "${IMG}" +echo "pushed ${IMG}" diff --git a/examples/inference/tensorrt-llm/nccl-ep-efa/setup/env_vars.example b/examples/inference/tensorrt-llm/nccl-ep-efa/setup/env_vars.example new file mode 100644 index 000000000..22eeffc43 --- /dev/null +++ b/examples/inference/tensorrt-llm/nccl-ep-efa/setup/env_vars.example @@ -0,0 +1,27 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. SPDX-License-Identifier: MIT-0 +# Copy to setup/env_vars (gitignored) and edit. Sourced by setup/build-push.sh; `source setup/env_vars` +# before running recipe/*.sh from a shell (in-pod, the kubernetes manifest env supplies these). +export REGISTRY=".dkr.ecr..amazonaws.com" # YOUR ECR — no registry is hardcoded anywhere +export IMAGE_NAME="trtllm-nccl-ep-efa" +# Immutable tag — never "latest": with the manifest's imagePullPolicy: IfNotPresent, a node +# that cached "latest" silently keeps running the OLD image after a rebuild-and-push. +export IMAGE_TAG="v1-20260825" +export AWS_REGION="us-east-2" +# Model: any MoE whose routed-expert count divides by the EP size (serve.sh preflights this). +# Default is public + ungated. The measured E2E used exactly this model at EP8. +export SERVE_MODEL="Qwen/Qwen3-30B-A3B" +# Pin the model revision (commit SHA or tag) alongside the name: --trust_remote_code runs +# repo Python in a privileged root container, so what executes should be an immutable +# artifact, not "whatever main points at today". Empty = HF default branch. Look up a SHA on +# the model's HF page ("Files and versions") and paste it here for a reproducible serve. +export SERVE_REVISION="" +export SERVE_TP="8" # single-node: tp_size == ep_size == GPUs on the node +export SERVE_EP="8" +# Algorithm selection is meaningful ONLY on an image built with APPLY_HT_FLAT_PATCH=1 +# (NVIDIA/TensorRT-LLM PR#17715). The unpatched upstream default is LOW_LATENCY+RANK_MAJOR, +# which is the measured-clean baseline on this substrate. serve.sh refuses a non-default +# algorithm on an unpatched image rather than silently ignoring it. +export TRTLLM_NCCL_EP_ALGO="LOW_LATENCY" +# Build the opt-in patched image: APPLY_HT_FLAT_PATCH=1 setup/build-push.sh +# (:- default so a value pre-set on the command line survives this file being sourced after it) +export APPLY_HT_FLAT_PATCH="${APPLY_HT_FLAT_PATCH:-0}" diff --git a/examples/inference/tensorrt-llm/nccl-ep-efa/setup_trtllm_nccl_ep_efa.sh b/examples/inference/tensorrt-llm/nccl-ep-efa/setup_trtllm_nccl_ep_efa.sh new file mode 100644 index 000000000..1c155d942 --- /dev/null +++ b/examples/inference/tensorrt-llm/nccl-ep-efa/setup_trtllm_nccl_ep_efa.sh @@ -0,0 +1,48 @@ +#!/usr/bin/env bash +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. SPDX-License-Identifier: MIT-0 +# +# setup_trtllm_nccl_ep_efa.sh — build aws-ofi-nccl's GIN (CPU-proxy) plugin for the +# TensorRT-LLM NcclEP backend on EFA. `nccl.ep` (nccl4py) drives its network traffic through +# NCCL, and on EFA that means aws-ofi-nccl with GIN compiled in; the EFA installer's stock +# plugin does not carry it, hence this source build. +# +# Distinct from the NVSHMEM-path setup_deepep_efa.sh (vendor-synced via +# .github/workflows/deepep-vendor-sync.yml) and from the vLLM sample's +# setup_deepep_v2_efa.sh — this script builds NO DeepEP at all (TRT-LLM's EP backend is +# nccl.ep, not the deep_ep python package) and is intentionally NOT vendor-synced. +# +# Runs inside the Docker build (no GPU needed). +set -euo pipefail + +# ---- pin (released tag; no 'latest') ---- +# v1.21.1 is the released tag that carries the CPU-proxy GIN op-tables this sample uses +# (src/rdma/gin/nccl_ofi_gin_api.cpp exports ncclGinPlugin_v11 + _v13; only _v14 is +# EFA-GDA-specific, which we do not use). It is the same tag the sibling +# micro-benchmarks/expert-parallelism/deepep-v2-benchmark/deepep.Dockerfile builds from +# source, so it is known-good in this repo. The plugin vendors its own GIN headers +# (3rd-party/nccl/cuda/include/nccl/gin_v13.h), so its GIN interface is not coupled to the +# pip NCCL headers — which is why no --with-nccl-headers flag is needed (and why that flag, +# not being an AC_ARG_WITH this project defines, was silently ignored before). +AWS_OFI_NCCL_REPO="${AWS_OFI_NCCL_REPO:-https://github.com/aws/aws-ofi-nccl.git}" +AWS_OFI_NCCL_REF="${AWS_OFI_NCCL_REF:-v1.21.1}" + +echo "== aws-ofi-nccl GIN @ ${AWS_OFI_NCCL_REF} ==" +git clone --depth 1 --branch "${AWS_OFI_NCCL_REF}" "${AWS_OFI_NCCL_REPO}" /opt/aws-ofi-nccl-src +cd /opt/aws-ofi-nccl-src +git rev-parse HEAD > /opt/aws-ofi-nccl.effective.sha +./autogen.sh +# Released v1.21.1 already attempts gdr_pin_buffer_v2 with GDR_PIN_FLAG_FORCE_PCIE and falls +# back to flags=0 on failure — so the forced-PCIe attempt is the default and needs no env +# override. The gdrdrv-2.4 kernel-module fallback that the old dev-line pin carried is a host +# precondition instead: gdrdrv >= 2.5 on the compute nodes (see README Prerequisites). +./configure --prefix=/opt/aws-ofi-nccl --with-libfabric=/opt/amazon/efa --with-cuda=/usr/local/cuda \ + --enable-cudart-dynamic --enable-platform-aws +make -C src -j"$(nproc)"; make -C src install +test -f /opt/aws-ofi-nccl/lib/libnccl-net-ofi.so +[ "$(nm -D /opt/aws-ofi-nccl/lib/libnccl-net-ofi.so | grep -c ncclGinPlugin)" -ge 1 ] # GIN symbol present (fail-loud) +# GIN needs gdrcopy COMPILED IN (gdrapi.h at configure time) — a gdrapi-less build carries +# this exact runtime-warn string and fails nccl_ofi_gin_init at serve. Assert absence. +[ "$(strings /opt/aws-ofi-nccl/lib/libnccl-net-ofi.so | grep -c 'GDRCopy support not available at compile time')" -eq 0 ] +ldconfig +cd /; rm -rf /opt/aws-ofi-nccl-src +echo "== setup_trtllm_nccl_ep_efa.sh complete: GIN plugin at /opt/aws-ofi-nccl/lib/libnccl-net-ofi.so =="