diff --git a/MODELS.md b/MODELS.md index 5d8185c86a..33c6fcad4d 100644 --- a/MODELS.md +++ b/MODELS.md @@ -60,7 +60,7 @@ Rationale: `dsv4` carries the largest single-turn footprint in the repository. 4 |---|---|---| | Agentic coding | Long Context, Multi Turn Realistic traffic trace replay with sub agents | Active. This is the trace-replay agentic-coding benchmark (see [`benchmarks/single_node/agentic/`](benchmarks/single_node/agentic/)). Going forward, new models will likely be onboarded with agentic coding only, and **with speculative decoding enabled only**. The non-spec-decode arm is not run or published (see [Deprecation Notice](#deprecation-notice)). | | Single-turn 8k1k | 8192 / 1024 | Active. This is the primary fixed-sequence-length scenario. | -| Single-turn 1k1k | 1024 / 1024 | **Deprecated for all models** since 2026-07-17 ([#2263](https://github.com/SemiAnalysisAI/InferenceX/pull/2263)), to save GPU cluster time for higher-priority real-world agentic-coding benchmarks and new frontier models. Archived configs live in [`configs/deprecated/`](configs/deprecated/). | +| Single-turn 1k1k | 1024 / 1024 | Deprecated since 2026-07-17 ([#2263](https://github.com/SemiAnalysisAI/InferenceX/pull/2263)), to save GPU cluster time for higher-priority real-world agentic-coding benchmarks and new frontier models. Archived configs live in [`configs/deprecated/`](configs/deprecated/). The GLM-5.1 B200 TileRT point added later in [#2533](https://github.com/SemiAnalysisAI/InferenceX/pull/2533) remains active. | | Single-turn 1k8k | 1024 / 8192 | **Deprecated for all models** since 2026-03-27 ([#911](https://github.com/SemiAnalysisAI/InferenceX/pull/911)), to save GPU cluster time for higher-priority real-world agentic-coding benchmarks and new frontier models. Configs were removed, not archived. | ## AgentX Guidelines @@ -157,7 +157,7 @@ Other offloading tiers, including NVMe KV cache offloading, are outside the init | MiniMax-M3 | `minimaxm3` | 2026-06-12 ([#1724](https://github.com/SemiAnalysisAI/InferenceX/pull/1724)) | Agentic coding | Single-turn 1k1k, Single-turn 8k1k (removed 2026-08-04, [#2493](https://github.com/SemiAnalysisAI/InferenceX/pull/2493)) | | DeepSeek-V4.1-Flash | `dsv41flash` | Pending | Agentic coding on MI355X (draft; GPU validation pending) | — | | DeepSeek-V4-Pro | `dsv4` | 2026-04-24 ([#1130](https://github.com/SemiAnalysisAI/InferenceX/pull/1130)) | Single-turn 8k1k (last day 2026-09-08, see [Deprecation Notice](#deprecation-notice)), Agentic coding (the non-MTP arm still runs while the MTP-only transition remains pending, as explained in the Deprecation Notice) | Single-turn 1k1k | -| GLM-5 / GLM-5.1 | `glm5`, `glm5.1` | 2026-03-06 ([#762](https://github.com/SemiAnalysisAI/InferenceX/pull/762)), with GLM-5.1 added 2026-04-21 ([#1098](https://github.com/SemiAnalysisAI/InferenceX/pull/1098)) | None (retired 2026-07-18, [#2276](https://github.com/SemiAnalysisAI/InferenceX/pull/2276)) | Single-turn 1k1k, Single-turn 1k8k (GLM-5 only), Single-turn 8k1k | +| GLM-5 / GLM-5.1 | `glm5`, `glm5.1` | 2026-03-06 ([#762](https://github.com/SemiAnalysisAI/InferenceX/pull/762)), with GLM-5.1 added 2026-04-21 ([#1098](https://github.com/SemiAnalysisAI/InferenceX/pull/1098)) | GLM-5.1 B200 TileRT only: 1k1k and 8k1k added 2026-08-09 ([#2533](https://github.com/SemiAnalysisAI/InferenceX/pull/2533)); Agentic coding added in [#2650](https://github.com/SemiAnalysisAI/InferenceX/pull/2650) | The earlier GLM-5 / GLM-5.1 recipes were retired 2026-07-18 ([#2276](https://github.com/SemiAnalysisAI/InferenceX/pull/2276)) | | MiniMax-M2.5/2.7 | `minimaxm2.5` | 2026-02-18 ([#755](https://github.com/SemiAnalysisAI/InferenceX/pull/755)) | None (retired 2026-06-20, [#1874](https://github.com/SemiAnalysisAI/InferenceX/pull/1874)) | Single-turn 1k1k, Single-turn 1k8k, Single-turn 8k1k | | Kimi-K2.5/2.6/2.7-Code | `kimik2.5` | 2026-02-17 ([#734](https://github.com/SemiAnalysisAI/InferenceX/pull/734)) | None (fully retired 2026-08-07, [#2527](https://github.com/SemiAnalysisAI/InferenceX/pull/2527)) | Single-turn 1k1k, Single-turn 1k8k, Agentic coding (removed 2026-08-04, [#2493](https://github.com/SemiAnalysisAI/InferenceX/pull/2493)), Single-turn 8k1k (removed 2026-08-07, [#2527](https://github.com/SemiAnalysisAI/InferenceX/pull/2527)) | | Qwen3.5-397B-A17B | `qwen3.5` | 2026-02-16 ([#704](https://github.com/SemiAnalysisAI/InferenceX/pull/704)) | Single-turn 8k1k and Agentic coding, both limited to fp8/fp4 | Single-turn 1k1k, Single-turn 1k8k, all bf16 recipes (removed 2026-08-04, [#2493](https://github.com/SemiAnalysisAI/InferenceX/pull/2493)) | diff --git a/MODELS_zh.md b/MODELS_zh.md index 29e1b55404..f2c3569b02 100644 --- a/MODELS_zh.md +++ b/MODELS_zh.md @@ -60,7 +60,7 @@ InferenceX-e2e 运行在数量固定且有限的 GPU 资源池上,并由一支 |---|---|---| | 智能体编码(agentic coding) | 长上下文、多轮真实流量的轨迹回放,含子智能体(sub agents) | 启用。此场景采用基于轨迹回放的智能体编码基准测试(见 [`benchmarks/single_node/agentic/`](benchmarks/single_node/agentic/))。今后新模型预计将仅以智能体编码场景接入,且**仅在启用投机解码的条件下**运行。非投机解码分支不运行也不发布(见[弃用公告](#弃用公告))。 | | 单轮 8k1k | 8192 / 1024 | 启用。当前主要的固定序列长度(fixed-seq-len)场景。 | -| 单轮 1k1k | 1024 / 1024 | **对所有模型均已弃用**,自 2026-07-17 起([#2263](https://github.com/SemiAnalysisAI/InferenceX/pull/2263)),以便将 GPU 集群时间留给优先级更高的真实场景智能体编码基准测试与新的前沿模型。归档配置位于 [`configs/deprecated/`](configs/deprecated/)。 | +| 单轮 1k1k | 1024 / 1024 | 自 2026-07-17 起弃用([#2263](https://github.com/SemiAnalysisAI/InferenceX/pull/2263)),以便将 GPU 集群时间留给优先级更高的真实场景智能体编码基准测试与新的前沿模型。归档配置位于 [`configs/deprecated/`](configs/deprecated/)。后续由 [#2533](https://github.com/SemiAnalysisAI/InferenceX/pull/2533) 加入的 GLM-5.1 B200 TileRT 测试点仍启用。 | | 单轮 1k8k | 1024 / 8192 | **对所有模型均已弃用**,自 2026-03-27 起([#911](https://github.com/SemiAnalysisAI/InferenceX/pull/911)),以便将 GPU 集群时间留给优先级更高的真实场景智能体编码基准测试与新的前沿模型。相关配置已删除,未归档。 | ## AgentX 指南 @@ -157,7 +157,7 @@ InferenceX 支持 SGLang 和 vLLM 双方的维护者,并响应 AI 实验室和 | MiniMax-M3 | `minimaxm3` | 2026-06-12([#1724](https://github.com/SemiAnalysisAI/InferenceX/pull/1724)) | 智能体编码 | 单轮 1k1k、单轮 8k1k(2026-08-04 移除,[#2493](https://github.com/SemiAnalysisAI/InferenceX/pull/2493)) | | DeepSeek-V4.1-Flash | `dsv41flash` | 待验证 | MI355X 上的 Agentic coding(草案;等待 GPU 验证) | — | | DeepSeek-V4-Pro | `dsv4` | 2026-04-24([#1130](https://github.com/SemiAnalysisAI/InferenceX/pull/1130)) | 单轮 8k1k、智能体编码(非 MTP 分支仍在运行,「仅 MTP」转换仍待执行,见弃用公告) | 单轮 1k1k | -| GLM-5 / GLM-5.1 | `glm5`、`glm5.1` | 2026-03-06([#762](https://github.com/SemiAnalysisAI/InferenceX/pull/762)),GLM-5.1 于 2026-04-21 加入([#1098](https://github.com/SemiAnalysisAI/InferenceX/pull/1098)) | 无(2026-07-18 退役,[#2276](https://github.com/SemiAnalysisAI/InferenceX/pull/2276)) | 单轮 1k1k、单轮 1k8k(仅 GLM-5)、单轮 8k1k | +| GLM-5 / GLM-5.1 | `glm5`、`glm5.1` | 2026-03-06([#762](https://github.com/SemiAnalysisAI/InferenceX/pull/762)),GLM-5.1 于 2026-04-21 加入([#1098](https://github.com/SemiAnalysisAI/InferenceX/pull/1098)) | 仅 GLM-5.1 B200 TileRT:1k1k 和 8k1k 于 2026-08-09 加入([#2533](https://github.com/SemiAnalysisAI/InferenceX/pull/2533));智能体编码由 [#2650](https://github.com/SemiAnalysisAI/InferenceX/pull/2650) 加入 | 此前的 GLM-5 / GLM-5.1 配方于 2026-07-18 退役([#2276](https://github.com/SemiAnalysisAI/InferenceX/pull/2276)) | | MiniMax-M2.5/2.7 | `minimaxm2.5` | 2026-02-18([#755](https://github.com/SemiAnalysisAI/InferenceX/pull/755)) | 无(2026-06-20 退役,[#1874](https://github.com/SemiAnalysisAI/InferenceX/pull/1874)) | 单轮 1k1k、单轮 1k8k、单轮 8k1k | | Kimi-K2.5/2.6/2.7-Code | `kimik2.5` | 2026-02-17([#734](https://github.com/SemiAnalysisAI/InferenceX/pull/734)) | 无(2026-08-07 完全退役,[#2527](https://github.com/SemiAnalysisAI/InferenceX/pull/2527)) | 单轮 1k1k、单轮 1k8k、智能体编码(2026-08-04 移除,[#2493](https://github.com/SemiAnalysisAI/InferenceX/pull/2493))、单轮 8k1k(2026-08-07 移除,[#2527](https://github.com/SemiAnalysisAI/InferenceX/pull/2527)) | | Qwen3.5-397B-A17B | `qwen3.5` | 2026-02-16([#704](https://github.com/SemiAnalysisAI/InferenceX/pull/704)) | 单轮 8k1k 与智能体编码,二者均仅限 fp8/fp4 | 单轮 1k1k、单轮 1k8k、全部 bf16 配方(2026-08-04 移除,[#2493](https://github.com/SemiAnalysisAI/InferenceX/pull/2493)) | diff --git a/benchmarks/multi_node/tilert_utils/run_node.sh b/benchmarks/multi_node/tilert_utils/run_node.sh index 0b92099c15..8961b2f7c9 100755 --- a/benchmarks/multi_node/tilert_utils/run_node.sh +++ b/benchmarks/multi_node/tilert_utils/run_node.sh @@ -232,7 +232,7 @@ wait_for_tcp() { fi sleep 5 done - [[ $rc -eq 0 ]] && { exec 3>&- 2>/dev/null || true; echo "[wait_for_tcp] $host:$port ready"; } + [[ $rc -eq 0 ]] && echo "[wait_for_tcp] $host:$port ready" (( _xtrace )) && set -x return $rc } @@ -261,11 +261,11 @@ run_bench_and_eval() { || { rc=$?; echo "[bench] WARNING: conc=$conc failed/timed out (rc=$rc)"; } done fi - run_lm_eval || rc=$? + run_tilert_eval || rc=$? return $rc } -run_lm_eval() { +run_tilert_eval() { [[ "${RUN_EVAL}" = "true" ]] || return 0 if [[ -n "${EVAL_CONC:-}" ]]; then export EVAL_CONCURRENT_REQUESTS="$EVAL_CONC" @@ -273,8 +273,13 @@ run_lm_eval() { export EVAL_CONCURRENT_REQUESTS="$(tr ' ' '\n' <<< "$CONC_LIST" | sort -n | tail -1)" fi export CONC="$EVAL_CONCURRENT_REQUESTS" - run_eval --port "$ROUTER_PORT" - append_lm_eval_summary + local eval_rc=0 stage_rc=0 + run_eval --port "$ROUTER_PORT" || eval_rc=$? + append_lm_eval_summary || stage_rc=$? + if [[ "$eval_rc" -ne 0 ]]; then + return "$eval_rc" + fi + return "$stage_rc" } run_agentic_replay() { diff --git a/benchmarks/multi_node/tilert_utils/submit.sh b/benchmarks/multi_node/tilert_utils/submit.sh index e29ae13988..302cdcbe83 100755 --- a/benchmarks/multi_node/tilert_utils/submit.sh +++ b/benchmarks/multi_node/tilert_utils/submit.sh @@ -113,13 +113,23 @@ trap 'exit 130' INT import_image() { local image_ref="$1" squash_file="$2" host="$3" + local enroot_ref="${image_ref#docker://}" + local registry="${enroot_ref%%/*}" + # Enroot needs '#' for an explicit registry; '/' alone targets Docker Hub. + if [[ "$enroot_ref" != *#* && "$enroot_ref" == */* && ( + "$registry" == *.* || "$registry" == *:* || "$registry" == localhost + ) ]]; then + enroot_ref="$registry#${enroot_ref#*/}" + fi local image_key; image_key=$(echo "$image_ref" | sed 's/[\/:@#]/_/g') local lock_file="$SQUASH_DIR/.locks/${image_key}.lock" mkdir -p "$SQUASH_DIR/.locks" srun --jobid="$JOB_ID" --nodelist="$host" --ntasks=1 bash -c " export ENROOT_CACHE_PATH=\$HOME/.cache/enroot; mkdir -p \$ENROOT_CACHE_PATH exec 9>\"$lock_file\"; flock -w 600 9 || exit 1 - unsquashfs -l \"$squash_file\" >/dev/null 2>&1 || enroot import -o \"$squash_file\" docker://$image_ref + unsquashfs -l \"$squash_file\" >/dev/null 2>&1 || { + rm -f \"$squash_file\" && enroot import -o \"$squash_file\" \"docker://$enroot_ref\" + } " } import_image "$DECODE_IMAGE" "$DECODE_SQUASH" "$DECODE_HOST" || exit 1 diff --git a/configs/nvidia-master.yaml b/configs/nvidia-master.yaml index b32f5e80ec..de21a8e3f8 100644 --- a/configs/nvidia-master.yaml +++ b/configs/nvidia-master.yaml @@ -10295,6 +10295,11 @@ glm5.1-fp8-b200-tilert: additional-settings: - "PREFILL_IMAGE=vllm/vllm-openai:v0.26.0" - "PREFILL_NODES=1" + - "SALLOC_TIME_LIMIT=45" + - "B200_SQUASH_DIR=/data/home/sa-shared/containers" + - "MODEL_PATH=/data/home/sa-shared/gharunners/hf-hub-cache/hub/models--zai-org--GLM-5.1-FP8/snapshots/f396cf805182f4ca10fa675e1a99815b3ca384db" + - "HF_HUB_CACHE_HOST_PATH=/data/home/sa-shared/gharunners/hf-hub-cache" + - "TILERT_WEIGHTS_DIR=/data/home/sa-shared/gharunners/tilert-cache/glm5.1-fp8-8shard" decode: num-worker: 1 tp: 8 @@ -10316,6 +10321,11 @@ glm5.1-fp8-b200-tilert: additional-settings: - "PREFILL_IMAGE=vllm/vllm-openai:v0.26.0" - "PREFILL_NODES=1" + - "SALLOC_TIME_LIMIT=90" + - "B200_SQUASH_DIR=/data/home/sa-shared/containers" + - "MODEL_PATH=/data/home/sa-shared/gharunners/hf-hub-cache/hub/models--zai-org--GLM-5.1-FP8/snapshots/f396cf805182f4ca10fa675e1a99815b3ca384db" + - "HF_HUB_CACHE_HOST_PATH=/data/home/sa-shared/gharunners/hf-hub-cache" + - "TILERT_WEIGHTS_DIR=/data/home/sa-shared/gharunners/tilert-cache/glm5.1-fp8-8shard" decode: num-worker: 1 tp: 8 diff --git a/docs/configuration-procedures.md b/docs/configuration-procedures.md index 1204fb08e4..805254c4ff 100644 --- a/docs/configuration-procedures.md +++ b/docs/configuration-procedures.md @@ -124,6 +124,12 @@ directory is not a completion signal. ## Native TileRT power +TileRT's shared importer preserves Docker Hub image names and converts explicit registries such as `ghcr.io/team/image:tag` to Enroot's `docker://ghcr.io#team/image:tag` syntax. Existing `#` references are preserved. Valid cached squash images are reused without importing; a cache hit does not validate the registry import path. Invalid cached images are removed under the import lock before retrying the import. + +The GLM-5.1 B200 Nscale 1k1k and 8k1k recipes select the prepared shared checkpoint, converted TileRT weights and squash cache, with allocation limits of 45 minutes for 1k1k and 90 minutes for 8k1k, including its full GSM8K eval. Since C1 is below automatic eval selection, use the PR `all-evals` label alongside `full-sweep-fail-fast` for full qualification. TileRT was added after the general GLM-5.1 retirement in [#2533](https://github.com/SemiAnalysisAI/InferenceX/pull/2533); [MODELS.md](../MODELS.md) records this retained scope. Changes still require the normal PR sweep, applicable quality evidence, sign-off and reuse before publication. + +TileRT's eval wrapper calls the shared `run_eval` dispatcher without overriding its `run_lm_eval` client. It stages available artifacts after evaluation and preserves failures from either evaluation or staging. TCP readiness probes keep their socket inside a subshell and preserve the caller's diagnostic streams. + For GLM-5.1 on B200 Nscale, `MODEL_PATH` can select an existing shared checkpoint instead of the default `/scratch/models/GLM-5.1-FP8`. When it selects an HF snapshot, also set `HF_HUB_CACHE_HOST_PATH` to the existing cache root; TileRT mounts that root at the same absolute path so snapshot links to sibling blobs remain readable. Keep `TILERT_WEIGHTS_DIR` pointed at the separately converted decode weights. Only fixed 8192/1024 `glm5.1-fp8-b200-tilert` requires native power. TileRT runs inside its returned `salloc` allocation, retains both role exit codes and drains collectors before staging audits. Exactly one physical node per role is supported. Other sequence lengths, AgentX and eval-only do not enable this collector. Hardware qualification and publication remain pending. diff --git a/docs/configuration-procedures_zh.md b/docs/configuration-procedures_zh.md index d7a48661df..900b7a8b94 100644 --- a/docs/configuration-procedures_zh.md +++ b/docs/configuration-procedures_zh.md @@ -122,6 +122,12 @@ B300 DSXE 的 Kimi-K3 AgentX 路径在 `/scratch/models` 下挂载预置目标 ## TileRT 原生功耗 +TileRT 的共享导入器保留 Docker Hub 镜像名称,并将 `ghcr.io/team/image:tag` 等显式仓库地址转换为 Enroot 的 `docker://ghcr.io#team/image:tag` 格式。已有的 `#` 地址保持不变。有效的缓存 squash 镜像会直接复用;命中缓存不能证明仓库导入路径有效。无效的缓存镜像会在持有导入锁时删除,再重新导入。 + +GLM-5.1 B200 Nscale 1k1k 和 8k1k 配方使用已准备的共享 checkpoint、TileRT 转换权重和 squash 缓存,1k1k 的分配时限为 45 分钟,8k1k 为 90 分钟,以容纳完整 GSM8K eval。C1 低于自动 eval 选择门槛,完整资格验证应同时使用 PR 标签 `all-evals` 和 `full-sweep-fail-fast`。TileRT 在 GLM-5.1 一般退役之后由 [#2533](https://github.com/SemiAnalysisAI/InferenceX/pull/2533) 加入;[MODELS_zh.md](../MODELS_zh.md) 记录了这部分保留范围。相关改动仍须完成正常 PR sweep、适用质量验证、签核和复用,才能发布。 + +TileRT 的 eval 封装调用共享 `run_eval` 分发器,不覆盖其中的 `run_lm_eval` 客户端。评测后保存可用产物,并保留评测或产物保存阶段的失败状态。TCP 就绪探测只在子 shell 中使用 socket,不改变调用方的诊断输出流。 + B200 Nscale 的 GLM-5.1 可用 `MODEL_PATH` 指定已有共享权重,覆盖默认的 `/scratch/models/GLM-5.1-FP8`。若指定 HF snapshot,还需把 `HF_HUB_CACHE_HOST_PATH` 设为现有缓存根目录;TileRT 按相同绝对路径挂载整个缓存,使 snapshot 指向同级 blobs 的软链接可读。`TILERT_WEIGHTS_DIR` 仍指向单独转换的 decode 权重。 仅固定 8192/1024 的 `glm5.1-fp8-b200-tilert` 要求原生功耗。TileRT 在 `salloc` 返回的分配内运行,保留两个角色的退出码,并在保存审计数据前等待采集器排空。每个角色仅支持一个物理节点。其他序列长度、AgentX 和 eval-only 不启用此采集器。硬件资格验证与发布仍待完成。 diff --git a/perf-changelog.yaml b/perf-changelog.yaml index 82e6ef4bf9..c2a59cb1b5 100644 --- a/perf-changelog.yaml +++ b/perf-changelog.yaml @@ -7581,3 +7581,16 @@ - "DP-attention serving adds --enable-dp-lm-head, which SGLang requires for DSpark under DP attention, and TP-only serving adds --prefill-decode-interval 10 to fix scheduler issue." - "Use --cuda-graph-max-bs-decode instead of --cuda-graph-max-bs: the 20260911 image split the flag per phase and ships no legacy alias, so the old spelling aborts argument parsing as an ambiguous prefix." pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3001 + +- config-keys: + - glm5.1-fp8-b200-tilert + scenario-type: + - fixed-seq-len + description: + - "Preserve explicit registries during native TileRT Enroot imports and select prepared shared weights and squash images for both C1 recipes." + - "原生 TileRT 导入镜像时保留显式仓库地址,并让两个 C1 配方使用已准备的共享权重与 squash 镜像。" + - "Allow 90 minutes for the 8k1k recipe and its full GSM8K evaluation." + - "为 8k1k 配方及其完整 GSM8K 评测预留 90 分钟。" + - "Avoid recursive TileRT eval dispatch and preserve evaluation failures after artifact staging." + - "避免 TileRT 评测分发无限递归,并在保存产物后保留评测失败状态。" + pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3067 diff --git a/runners/test_tilert_power_lifecycle.py b/runners/test_tilert_power_lifecycle.py index 973e7595ed..c208965e93 100644 --- a/runners/test_tilert_power_lifecycle.py +++ b/runners/test_tilert_power_lifecycle.py @@ -10,6 +10,64 @@ ROOT = Path(__file__).resolve().parents[1] +@pytest.mark.parametrize(('image', 'expected_uri', 'cached'), [ + ('ghcr.io/tile-ai/tilert:0.1.5', 'docker://ghcr.io#tile-ai/tilert:0.1.5', None), + ('vllm/vllm-openai:v0.26.0', 'docker://vllm/vllm-openai:v0.26.0', None), + ('ghcr.io#tile-ai/tilert:0.1.5', 'docker://ghcr.io#tile-ai/tilert:0.1.5', None), + ('docker://ghcr.io#tile-ai/tilert:0.1.5', 'docker://ghcr.io#tile-ai/tilert:0.1.5', None), + ('registry.example:5000/team/image:tag', 'docker://registry.example:5000#team/image:tag', None), + ('ghcr.io/tile-ai/tilert:0.1.5', None, 'prepared image'), + pytest.param('ghcr.io/tile-ai/tilert:0.1.5', 'docker://ghcr.io#tile-ai/tilert:0.1.5', + 'partial image', id='partial-cache'), +]) +def test_tilert_import_uses_registry_and_reuses_squash(tmp_path, image, expected_uri, cached): + bindir, squash = tmp_path / 'bin', tmp_path / 'squash' + bindir.mkdir() + squash.mkdir() + image_file = squash / (image.translate(str.maketrans('/:@#', '____')) + '.sqsh') + if cached is not None: + image_file.write_text(cached) + scripts = { + 'scontrol': '#!/bin/sh\nprintf "node-a\\nnode-b\\n"\n', + 'flock': '#!/bin/sh\nexit 0\n', + 'unsquashfs': '#!/bin/sh\ngrep -qxE "prepared image|imported image" "$2"\n', + 'enroot': '#!' + sys.executable + '\n' + ''' +import json, os, sys +assert sys.argv[1:3] == ['import', '-o'] +with open(os.environ['IMPORT_RECEIPT'], 'a') as receipt: + receipt.write(json.dumps(sys.argv[4:]) + '\\n') +with open(sys.argv[3], 'x') as image: + image.write('imported image') +''', + 'srun': '#!' + sys.executable + '\n' + ''' +import os, subprocess, sys +args = sys.argv[1:] +if any(arg.startswith('--container-image=') for arg in args): + sys.exit(0) +command = args[next(i for i, arg in enumerate(args) if not arg.startswith('--')):] +os.execvpe(command[0], command, os.environ) +''', + } + for name, script in scripts.items(): + path = bindir / name + path.write_text(script) + path.chmod(0o755) + receipt = tmp_path / 'imports.jsonl' + env = {**os.environ, 'PATH': str(bindir) + os.pathsep + os.environ['PATH'], + 'GITHUB_WORKSPACE': str(tmp_path), 'B200_SQUASH_DIR': str(squash), + 'IMAGE': image, 'DECODE_IMAGE': image, 'PREFILL_IMAGE': image, + 'MODEL_PATH': str(tmp_path), 'TILERT_WEIGHTS_DIR': str(tmp_path / 'weights'), + 'TILERT_IN_ALLOCATION': '1', 'SLURM_JOB_ID': '123', + 'SLURM_JOB_NODELIST': 'node-[a-b]', 'ISL': '1024', 'OSL': '1024', + 'REQUIRE_POWER': '0', 'IMPORT_RECEIPT': str(receipt), 'HOME': str(tmp_path)} + result = subprocess.run(['bash', str(ROOT / 'benchmarks/multi_node/tilert_utils/submit.sh')], + env=env, capture_output=True, text=True, timeout=10) + assert result.returncode == 0, result.stderr + imports = [json.loads(line) for line in receipt.read_text().splitlines()] if receipt.exists() else [] + assert imports == ([] if expected_uri is None else [[expected_uri]]) + assert image_file.read_text() == ('prepared image' if expected_uri is None else 'imported image') + + @pytest.mark.parametrize('prepared_path', ['', '/shared/hf/hub/snapshots/revision']) def test_b200_tilert_preserves_prepared_model_path(prepared_path): env = {**os.environ, 'MODEL_PREFIX': 'glm5.1', 'PRECISION': 'fp8', @@ -156,3 +214,59 @@ def test_tilert_node_publishes_complete_owned_sentinel(tmp_path, node_rc): capture_output=True, text=True, timeout=5) assert result.returncode == node_rc, result.stderr assert (tmp_path / 'done').read_text().strip() == str(node_rc) + + +@pytest.mark.parametrize(('eval_rc', 'stage_rc'), [(0, 0), (7, 0), (0, 9), (7, 9)]) +def test_tilert_eval_dispatches_shared_client_and_preserves_failure(tmp_path, eval_rc, stage_rc): + source = (ROOT / 'benchmarks/multi_node/tilert_utils/run_node.sh').read_text() + functions = source[source.index('run_bench_and_eval() {'):source.index('run_agentic_replay() {')] + command = ''' +source "$1/benchmarks/benchmark_lib.sh" +FUNCNEST=40 +wait_for_server_ready() { return 0; } +run_server_client() { printf '%s\\n' "$@" > client-args; return "$CLIENT_RC"; } +append_lm_eval_summary() { touch staged; return "$STAGE_RC"; } +''' + functions + '\nrun_bench_and_eval\n' + env = {**os.environ, 'RUN_EVAL': 'true', 'EVAL_ONLY': 'true', + 'POWERX_NATIVE_ENABLED': '0', 'ROUTER_PORT': '9876', 'ROUTER_PID': '1', + 'BENCHMARK_LOGS_DIR': str(tmp_path), 'CONC_LIST': '1', 'EVAL_CONC': '2', + 'MODEL_NAME': 'test-model', 'MODEL': 'test-model', 'EVAL_MAX_MODEL_LEN': '9472', + 'EVAL_FRAMEWORK': 'lm-eval', 'EVAL_SUITE': '', 'EVAL_TASKS_DIR': 'gsm8k', + 'EVAL_RESULT_DIR': str(tmp_path / 'eval'), 'OPENAI_API_KEY': 'EMPTY', + 'INFERENCEX_LM_EVAL_RUNTIME_READY': 'true', 'IS_AGENTIC': '0', + 'SCENARIO_TYPE': 'single_turn', 'PYTHONPYCACHEPREFIX': str(tmp_path / 'pycache'), + 'CLIENT_RC': str(eval_rc), 'STAGE_RC': str(stage_rc)} + result = subprocess.run(['bash', '-c', command, 'bash', str(ROOT)], + env=env, cwd=tmp_path, capture_output=True, text=True, timeout=5) + assert result.returncode == (eval_rc or stage_rc), result.stderr + result.stdout + args = (tmp_path / 'client-args').read_text().splitlines() + assert args[:3] == ['python3', '-m', 'lm_eval'] + model_args = args[args.index('--model_args') + 1].split(',') + assert 'base_url=http://0.0.0.0:9876/v1/chat/completions' in model_args + assert 'num_concurrent=2' in model_args + assert (tmp_path / 'staged').exists() + + +def test_tilert_tcp_wait_preserves_caller_streams(tmp_path): + import re + import socket + + source = (ROOT / 'benchmarks/multi_node/tilert_utils/run_node.sh').read_text() + function = re.search(r'^wait_for_tcp\(\) \{\n.*?^\}', source, + flags=re.MULTILINE | re.DOTALL).group() + with socket.socket() as server: + server.bind(('127.0.0.1', 0)) + server.listen(1) + command = function + ''' +exec 3>caller-fd +wait_for_tcp 127.0.0.1 "$1" 0 +rc=$? +printf 'eval diagnostic\\n' >&2 +printf 'caller stream\\n' >&3 +exit "$rc" +''' + result = subprocess.run(['bash', '-c', command, 'bash', str(server.getsockname()[1])], + cwd=tmp_path, capture_output=True, text=True, timeout=5) + assert result.returncode == 0, result.stderr + assert 'eval diagnostic' in result.stderr + assert (tmp_path / 'caller-fd').read_text() == 'caller stream\n'