Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/AGENT_OPERATIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -96,7 +96,7 @@ Multinode disaggregated results add `prefill_gpu_energy_j`, `decode_gpu_energy_j

Every power result — valid or invalid, single-node or multinode — carries `power_metric_schema_version`. Version 2 defines each unprefixed `joules_per_*` field as whole-deployment GPU-board energy over the named denominator; role-scoped energy uses the explicit `prefill_*` / `decode_*` keys. Rows without the field predate the whole-deployment switch and their unprefixed joules are not comparable across topologies.

For srt-slurm recipes, `telemetry: {provider: dcgm-power}` enables official energy collection. `runners/launch_gb200-nv.sh` and `runners/launch_gb300-nv.sh` are the source of truth for `POWER_SRT_SLURM_PIN`. CI derives `POWER_PRODUCER_SHA` from the launcher stamp. `utils/test_gb200_power_official_contract.py` and `utils/test_gb300_power_official_contract.py` enforce the recipe/launcher contract. Only `PRECISION=fp8` dcgm-power lanes are validated.
For srt-slurm recipes, `telemetry: {provider: dcgm-power}` enables official energy collection. `runners/launch_gb200-nv.sh` and `runners/launch_gb300-nv.sh` are the source of truth for `POWER_SRT_SLURM_PIN`. CI derives `POWER_PRODUCER_SHA` from the launcher stamp. `utils/test_gb200_power_official_contract.py` and `utils/test_gb300_power_official_contract.py` enforce the recipe/launcher contract. Eligible recipe-gated `dynamo-sglang` dcgm-power lanes are validated.

Power audit artifacts are named `power_audit_<result>` and contain `power_validation_<result>.json` for single-node runs or `power_validation_<result>_*.json` for multinode runs. They are uploaded even when validation fails.

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -154,3 +154,16 @@ benchmark:
concurrencies: "512"
req_rate: "inf"
use_chat_template: false

telemetry:
enabled: true
provider: dcgm-power
default_frequency: 1.0
storage_subdir: power
required: true
startup_timeout_seconds: 120
request_timeout_seconds: 2
collector_join_timeout_seconds: 10
dcgm_exporter:
container_image: dcgm-exporter
port: 9401
Original file line number Diff line number Diff line change
Expand Up @@ -115,3 +115,16 @@ benchmark:
concurrencies: "1"
req_rate: "inf"
use_chat_template: false

telemetry:
enabled: true
provider: dcgm-power
default_frequency: 1.0
storage_subdir: power
required: true
startup_timeout_seconds: 120
request_timeout_seconds: 2
collector_join_timeout_seconds: 10
dcgm_exporter:
container_image: dcgm-exporter
port: 9401
Original file line number Diff line number Diff line change
Expand Up @@ -154,3 +154,16 @@ benchmark:
concurrencies: "256"
req_rate: "inf"
use_chat_template: false

telemetry:
enabled: true
provider: dcgm-power
default_frequency: 1.0
storage_subdir: power
required: true
startup_timeout_seconds: 120
request_timeout_seconds: 2
collector_join_timeout_seconds: 10
dcgm_exporter:
container_image: dcgm-exporter
port: 9401
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ name: "disagg-gb300-10p1d-dep4-dep32-18-c2500"

model:
path: "deepseek-v4-pro"
container: "lmsysorg/sglang:nightly-dev-cu13-20260707-b4155233"
container: "lmsysorg/sglang:v0.5.14-cu130@sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3"
precision: "fp4"

dynamo:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ name: "disagg-gb300-12p1d-dep4-dep24-18-c3000"

model:
path: "deepseek-v4-pro"
container: "lmsysorg/sglang:nightly-dev-cu13-20260707-b4155233"
container: "lmsysorg/sglang:v0.5.14-cu130@sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3"
precision: "fp4"

dynamo:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ name: "disagg-gb300-14p1d-dep4-dep16-18-c8192"

model:
path: "deepseek-v4-pro"
container: "lmsysorg/sglang:nightly-dev-cu13-20260707-b4155233"
container: "lmsysorg/sglang:v0.5.14-cu130@sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3"
precision: "fp4"

dynamo:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ name: "disagg-gb300-15p1d-dep4-dep12-18-c12000"

model:
path: "deepseek-v4-pro"
container: "lmsysorg/sglang:nightly-dev-cu13-20260707-b4155233"
container: "lmsysorg/sglang:v0.5.14-cu130@sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3"
precision: "fp4"

dynamo:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -33,7 +33,7 @@ name: "disagg-gb300-1p1d-dep4-dep16-5-c1024"

model:
path: "deepseek-v4-pro"
container: "lmsysorg/sglang:nightly-dev-cu13-20260707-b4155233"
container: "lmsysorg/sglang:v0.5.14-cu130@sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3"
precision: "fp4"

dynamo:
Expand Down Expand Up @@ -177,3 +177,18 @@ benchmark:
req_rate: "inf"
use_chat_template: false
custom_tokenizer: "sa_bench_tokenizers.sglang_deepseek_v4.SGLangDeepseekV4Tokenizer"

telemetry:
enabled: true
provider: dcgm-power
default_frequency: 1.0
storage_subdir: power
required: true
startup_timeout_seconds: 120
request_timeout_seconds: 2
collector_join_timeout_seconds: 10
dcgm_exporter:
container_image: dcgm-exporter
# 9401 is already bound by the cluster-level exporter on im-gb300 nodes;
# use a port outside that range.
port: 19401
Original file line number Diff line number Diff line change
Expand Up @@ -33,7 +33,7 @@ name: "disagg-gb300-1p1d-tp4-tp4-2-c1"

model:
path: "deepseek-v4-pro"
container: "lmsysorg/sglang:nightly-dev-cu13-20260707-b4155233"
container: "lmsysorg/sglang:v0.5.14-cu130@sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3"
precision: "fp4"

# See ../1k1k/disagg-gb200-1p1d-dep8-tep8.yaml for the dynamo pin
Expand Down Expand Up @@ -160,3 +160,18 @@ benchmark:
req_rate: "inf"
use_chat_template: false
custom_tokenizer: "sa_bench_tokenizers.sglang_deepseek_v4.SGLangDeepseekV4Tokenizer"

telemetry:
enabled: true
provider: dcgm-power
default_frequency: 1.0
storage_subdir: power
required: true
startup_timeout_seconds: 120
request_timeout_seconds: 2
collector_join_timeout_seconds: 10
dcgm_exporter:
container_image: dcgm-exporter
# 9401 is already bound by the cluster-level exporter on im-gb300 nodes;
# use a port outside that range.
port: 19401
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ name: "disagg-gb300-8p1d-dep4-dep40-18-c2048"

model:
path: "deepseek-v4-pro"
container: "lmsysorg/sglang:nightly-dev-cu13-20260707-b4155233"
container: "lmsysorg/sglang:v0.5.14-cu130@sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3"
precision: "fp4"

dynamo:
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -173,3 +173,18 @@ benchmark:
concurrencies: "1x4x8x16x32x64x256"
req_rate: "inf"
random_range_ratio: 0.8

telemetry:
enabled: true
provider: dcgm-power
default_frequency: 1.0
storage_subdir: power
required: true
startup_timeout_seconds: 120
request_timeout_seconds: 2
collector_join_timeout_seconds: 10
dcgm_exporter:
container_image: dcgm-exporter
# 9401 is already bound by the cluster-level exporter on im-gb300 nodes;
# use a port outside that range.
port: 19401
2 changes: 1 addition & 1 deletion configs/nvidia-master.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -5875,7 +5875,7 @@ dsv4-fp4-gb300-dynamo-trt-mtp:
dp-attn: true

dsv4-fp4-gb300-dynamo-sglang:
image: lmsysorg/sglang:nightly-dev-cu13-20260707-b4155233
image: lmsysorg/sglang:v0.5.14-cu130@sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3
model: deepseek-ai/DeepSeek-V4-Pro
model-prefix: dsv4
runner: gb300
Expand Down
12 changes: 12 additions & 0 deletions perf-changelog.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -5926,6 +5926,18 @@
- "Rides on the NVFP4-V2 checkpoint switch from #2205"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2550

- config-keys:
- dsv4-fp4-gb200-dynamo-sglang
- dsv4-fp4-gb300-dynamo-sglang
- qwen3.5-fp4-gb300-dynamo-sglang
scenario-type:
- fixed-seq-len
description:
- "Extend official DCGM power collection to the recipe-gated FP4 GB200 and GB300 Dynamo-SGLang lanes."
- "Synchronize all DSV4 GB300 recipes on lmsysorg/sglang:v0.5.14-cu130@sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3."
- "Pin the srt-slurm power producer and record its exact checkout SHA for strict provenance validation."
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2507

- config-keys:
- qwen3.5-fp8-b300-sglang-agentic-power-ab
- qwen3.5-fp4-b300-sglang-agentic-power-ab
Expand Down
34 changes: 25 additions & 9 deletions runners/launch_gb200-nv.sh
Original file line number Diff line number Diff line change
Expand Up @@ -302,11 +302,12 @@ if [[ -n "$CONFIG_FILE" && -f "$_RECIPE_SRC" ]] && awk '
USES_DCGM_POWER=1
fi

# Note (wenyao): the producer pin descends from the fp8 srt-slurm lineage
# (cargo/maturin bootstrap); a non-fp8 power recipe would silently clone the
# wrong lineage, so fail fast instead.
if [[ "$USES_DCGM_POWER" == "1" && "$PRECISION" != "fp8" ]]; then
echo "Error: dcgm-power lanes are only validated for PRECISION=fp8, got: $PRECISION" >&2
# Note (wenyao): the producer pin follows the srt-slurm main lineage that the
# dynamo-sglang lanes run on (fp8 validated end-to-end, fp4 recipes
# parse-verified against the pin); other frameworks clone diverging refs
# (aflowers branch, sa-submission), so fail fast for them instead.
if [[ "$USES_DCGM_POWER" == "1" && "$FRAMEWORK" != "dynamo-sglang" ]]; then
echo "Error: dcgm-power lanes are only validated for FRAMEWORK=dynamo-sglang, got: $FRAMEWORK" >&2
Comment thread
edwingao28 marked this conversation as resolved.
exit 1
fi

Expand Down Expand Up @@ -475,10 +476,25 @@ elif [[ $FRAMEWORK == "dynamo-vllm" && $MODEL_PREFIX == "dsv4" ]]; then
mkdir -p recipes/vllm/deepseek-v4
cp -rT "$GITHUB_WORKSPACE/benchmarks/multi_node/srt-slurm-recipes/vllm/deepseek-v4" recipes/vllm/deepseek-v4
elif [[ $FRAMEWORK == "dynamo-sglang" && $MODEL_PREFIX == "dsv4" ]]; then
# Stay on NVIDIA/srt-slurm:main (default) — submission branch no
# longer needed; overlay our hand-rolled DSV4 sglang recipes onto it.
git clone https://github.com/NVIDIA/srt-slurm.git "$SRT_REPO_DIR"
cd "$SRT_REPO_DIR"
if [[ "$USES_DCGM_POWER" == "1" ]]; then
# Note (wenyao): on this cluster the DSV4-Pro checkpoint lives on the
# compute-node /mnt/numa1 NVMe (same staging the agentic path and the
# llm-d sweeps load from); the lustre alias target the shared dsv4
# block exports is not present here. Scoped to the power lane so the
# non-power lane keeps whatever the external-cluster staging expects.
export MODEL_PATH="/mnt/numa1/models/DeepSeek-V4-Pro"
git clone "$POWER_SRT_SLURM_URL" "$SRT_REPO_DIR"
cd "$SRT_REPO_DIR"
git checkout "$POWER_SRT_SLURM_PIN" || exit 1
# The power lane must run the exact pinned producer SHA, never a moving branch.
test "$(git rev-parse HEAD)" = "$POWER_SRT_SLURM_PIN" || { echo "Error: srt-slurm HEAD does not match POWER_SRT_SLURM_PIN=$POWER_SRT_SLURM_PIN" >&2; exit 1; }
git rev-parse HEAD > "$GITHUB_WORKSPACE/power-producer-sha.txt"
else
# Stay on NVIDIA/srt-slurm:main (default) — submission branch no
# longer needed; overlay our hand-rolled DSV4 sglang recipes onto it.
git clone https://github.com/NVIDIA/srt-slurm.git "$SRT_REPO_DIR"
cd "$SRT_REPO_DIR"
fi
mkdir -p recipes/sglang/deepseek-v4
cp -rT "$GITHUB_WORKSPACE/benchmarks/multi_node/srt-slurm-recipes/sglang/deepseek-v4" recipes/sglang/deepseek-v4
elif [[ $FRAMEWORK == "dynamo-sglang" && $MODEL_PREFIX == "glm5.1" ]]; then
Expand Down
45 changes: 30 additions & 15 deletions runners/launch_gb300-nv.sh
Original file line number Diff line number Diff line change
Expand Up @@ -155,11 +155,12 @@ if [[ -n "$CONFIG_FILE" && -f "$_RECIPE_SRC" ]] && awk '
USES_DCGM_POWER=1
fi

# Note (wenyao): the producer pin descends from the fp8 v1.0.25 lineage
# (cargo/maturin bootstrap); an fp4 power recipe would silently skip the
# sa-submission branch it needs, so fail fast instead.
if [[ "$USES_DCGM_POWER" == "1" && "$PRECISION" != "fp8" ]]; then
echo "Error: dcgm-power lanes are only validated for PRECISION=fp8, got: $PRECISION" >&2
# Note (wenyao): the producer pin follows the srt-slurm main lineage that the
# dynamo-sglang lanes run on (fp8 validated end-to-end, fp4 recipes
# parse-verified against the pin); other frameworks clone diverging refs
# (aflowers branch, sa-submission), so fail fast for them instead.
if [[ "$USES_DCGM_POWER" == "1" && "$FRAMEWORK" != "dynamo-sglang" ]]; then
Comment thread
edwingao28 marked this conversation as resolved.
echo "Error: dcgm-power lanes are only validated for FRAMEWORK=dynamo-sglang, got: $FRAMEWORK" >&2
exit 1
fi

Expand Down Expand Up @@ -271,14 +272,25 @@ elif [[ $FRAMEWORK == "dynamo-vllm" && $MODEL_PREFIX == "dsv4" ]]; then
cp -rT "$GITHUB_WORKSPACE/benchmarks/multi_node/srt-slurm-recipes/vllm/deepseek-v4" recipes/vllm/deepseek-v4
elif [[ $FRAMEWORK == "dynamo-sglang" && $MODEL_PREFIX == "dsv4" ]]; then
# Fixed-length DeepSeek-V4 recipes are version-controlled in this repository;
# overlay them onto the srt-slurm release that bootstraps cargo/maturin for
# the hash-pinned Dynamo source build before launch.
git clone https://github.com/NVIDIA/srt-slurm.git "$SRT_REPO_DIR"
cd "$SRT_REPO_DIR"
git checkout v1.0.25
mkdir -p recipes/sglang/deepseek-v4/8k1k
# overlay them onto the selected srt-slurm checkout. Power lanes use the
# exact producer pin and stamp it for strict result provenance; non-power
# lanes retain the v1.0.25 release that bootstraps cargo/maturin for the
# hash-pinned Dynamo source build before launch.
if [[ "$USES_DCGM_POWER" == "1" ]]; then
git clone "$POWER_SRT_SLURM_URL" "$SRT_REPO_DIR" || exit 1
cd "$SRT_REPO_DIR" || exit 1
git checkout "$POWER_SRT_SLURM_PIN" || exit 1
# The power lane must run the exact pinned producer SHA, never a moving branch.
test "$(git rev-parse HEAD)" = "$POWER_SRT_SLURM_PIN" || { echo "Error: srt-slurm HEAD does not match POWER_SRT_SLURM_PIN=$POWER_SRT_SLURM_PIN" >&2; exit 1; }
git rev-parse HEAD > "$GITHUB_WORKSPACE/power-producer-sha.txt"
else
git clone https://github.com/NVIDIA/srt-slurm.git "$SRT_REPO_DIR" || exit 1
cd "$SRT_REPO_DIR" || exit 1
git checkout v1.0.25 || exit 1
fi
mkdir -p recipes/sglang/deepseek-v4/8k1k || exit 1
cp -rT "$GITHUB_WORKSPACE/benchmarks/multi_node/srt-slurm-recipes/sglang/deepseek-v4/8k1k" \
recipes/sglang/deepseek-v4/8k1k
recipes/sglang/deepseek-v4/8k1k || exit 1
elif [[ $FRAMEWORK == "dynamo-sglang" && $MODEL_PREFIX == "glm5" ]]; then
git clone https://github.com/NVIDIA/srt-slurm.git "$SRT_REPO_DIR"
cd "$SRT_REPO_DIR"
Expand Down Expand Up @@ -472,16 +484,19 @@ inject_synthetic_acceptance "$CONFIG_PATH" "$FRAMEWORK" || exit 1
# - glm5.1, whose GLM-5.1-NVFP4 weights are prestaged on the compute-node
# /scratch/models, and
# - qwen3.5 fp8, whose weights are also on the compute-node /scratch/models
# and which runs on srt-slurm:v1.0.25 (the release that has the preflight;
# qwen3.5 fp4 runs on v1.0.29, which has none).
# and which runs on srt-slurm:v1.0.25 (the release that has the preflight),
# - qwen3.5 fp4 dynamo-trt, which runs on v1.0.29 without that preflight, and
# - the qwen3.5 fp4 and dsv4 sglang power lanes, which run the pinned
# producer (a main-lineage fork that has the preflight) against the same
# /scratch checkpoints.
# The engine still fails loudly at runtime if the path is genuinely missing on
# the compute node. Other fixed-seq-len recipes resolve model.path to a
# login-visible location, so keep the precheck enforced for them.
SRTCTL_APPLY_ARGS=(
-f "$CONFIG_FILE"
--tags "gb300,${MODEL_PREFIX},${PRECISION},${ISL}x${OSL},infmax-$(date +%Y%m%d)"
)
if [[ "$IS_AGENTIC" == "1" || "$MODEL_PREFIX" == "glm5.1" || ( "$MODEL_PREFIX" == "qwen3.5" && "$PRECISION" == "fp8" ) || ( "$MODEL_PREFIX" == "qwen3.5" && "$PRECISION" == "fp4" && "$FRAMEWORK" == "dynamo-trt" ) ]]; then
if [[ "$IS_AGENTIC" == "1" || "$MODEL_PREFIX" == "glm5.1" || ( "$MODEL_PREFIX" == "qwen3.5" && "$PRECISION" == "fp8" ) || ( "$MODEL_PREFIX" == "qwen3.5" && "$PRECISION" == "fp4" && ( "$FRAMEWORK" == "dynamo-trt" || "$USES_DCGM_POWER" == "1" ) ) || ( "$USES_DCGM_POWER" == "1" && "$MODEL_PREFIX" == "dsv4" && "$FRAMEWORK" == "dynamo-sglang" ) ]]; then
SRTCTL_APPLY_ARGS+=(--no-preflight)
fi
if [[ -n "$SRTCTL_SETUP_SCRIPT" ]]; then
Expand Down
34 changes: 28 additions & 6 deletions utils/test_gb200_power_official_contract.py
Original file line number Diff line number Diff line change
Expand Up @@ -61,7 +61,7 @@ def assert_recipe_driven_detection(launcher):
assert "t && /^ enabled: true$/ { e = 1 }" in launcher
assert "USES_DCGM_POWER=1" in launcher
assert (
'if [[ "$USES_DCGM_POWER" == "1" && "$PRECISION" != "fp8" ]]; then'
'if [[ "$USES_DCGM_POWER" == "1" && "$FRAMEWORK" != "dynamo-sglang" ]]; then'
in launcher
)
# Exporter provisioning, pinned clone, srtslurm.yaml injection, and audit
Expand Down Expand Up @@ -149,14 +149,36 @@ def test_pin_literal_lives_only_in_the_two_launchers():
]


def test_exactly_two_recipes_opt_into_dcgm_power():
# Every recipe that opts into dcgm-power, with the exporter port its cluster
# requires (im-gb300 nodes already bind 9401; gb200 keeps the default).
POWER_RECIPES = {
"benchmarks/multi_node/srt-slurm-recipes/sglang/deepseek-v4/8k1k/disagg-gb200-1p1d-dep8-dep16-6-c512.yaml": 9401,
"benchmarks/multi_node/srt-slurm-recipes/sglang/deepseek-v4/8k1k/disagg-gb200-1p1d-tp8-tp8-4-c1.yaml": 9401,
"benchmarks/multi_node/srt-slurm-recipes/sglang/deepseek-v4/8k1k/disagg-gb200-1p2d-dep8-dep16-10-c256.yaml": 9401,
"benchmarks/multi_node/srt-slurm-recipes/sglang/deepseek-v4/8k1k/disagg-gb300-1p1d-dep4-dep16-5-c1024.yaml": 19401,
"benchmarks/multi_node/srt-slurm-recipes/sglang/deepseek-v4/8k1k/disagg-gb300-1p1d-tp4-tp4-2-c1.yaml": 19401,
"benchmarks/multi_node/srt-slurm-recipes/sglang/qwen3.5/gb200-fp8/8k1k/1p1d-tp4-tp4.yaml": 9401,
"benchmarks/multi_node/srt-slurm-recipes/sglang/qwen3.5/gb300-fp4/8k1k/disagg/stp/8k1k_stp_lowlat_0.yaml": 19401,
"benchmarks/multi_node/srt-slurm-recipes/sglang/qwen3.5/gb300-fp8/8k1k/1p1d-tp4-tp4.yaml": 19401,
}


def test_exactly_the_declared_recipes_opt_into_dcgm_power():
hits = git_grep_lines(
"-lF", "provider: dcgm-power", "--", "benchmarks/multi_node/srt-slurm-recipes"
)
assert sorted(hits) == [
"benchmarks/multi_node/srt-slurm-recipes/sglang/qwen3.5/gb200-fp8/8k1k/1p1d-tp4-tp4.yaml",
"benchmarks/multi_node/srt-slurm-recipes/sglang/qwen3.5/gb300-fp8/8k1k/1p1d-tp4-tp4.yaml",
]
assert sorted(hits) == sorted(POWER_RECIPES)


def test_every_power_recipe_declares_the_same_telemetry_contract():
for path, port in POWER_RECIPES.items():
telemetry = yaml.safe_load((REPO_ROOT / path).read_text())["telemetry"]
assert telemetry["enabled"] is True, path
assert telemetry["provider"] == "dcgm-power", path
assert telemetry["default_frequency"] == 1.0, path
assert telemetry["required"] is True, path
assert telemetry["dcgm_exporter"]["container_image"] == "dcgm-exporter", path
assert telemetry["dcgm_exporter"]["port"] == port, path


def test_lane_and_pin_have_no_env_override_backdoor():
Expand Down
Loading