Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
37 changes: 23 additions & 14 deletions .github/workflows/nightly-bfcl.yml
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@ on:
workflow_dispatch:
inputs:
only:
description: "Run only this matrix leg (qwen3.8|gpt-oss|deepseek-v4.1|minimax-m2.7|kimi-k2.6|glm-5.2); empty = all"
description: "Run only this matrix leg (qwen3.8|gpt-oss|deepseek-v4.1|minimax-m2.7|kimi-k2.6|glm-5.3-flash); empty = all"
required: false
default: ""
model:
Expand Down Expand Up @@ -162,7 +162,7 @@ jobs:
# Concurrent TP=4 per arm (0-3 + 4-7).
# DeepSeek-V4.1-Flash is ~765GB on disk (552B backbone + 196B int8 Engram,
# all GPU-resident), so it outgrows a TP=4 half-node and runs SEQUENTIAL
# on the whole node like glm-5.2. SMG: deepseek_v41 tool/reasoning parsers
# on the whole node (the only such leg). SMG: deepseek_v41 tool/reasoning parsers
# + native renderer (#2524/#2525/#2527). vLLM: V4.1 is on main only
# (vllm-project/vllm#56208 + #56228, 2026-09-10), not in the 0.27.1 CI
# pin, so this leg installs a per-commit main wheel (vllm_commit below).
Expand Down Expand Up @@ -208,20 +208,29 @@ jobs:
"vllm_extra": "--trust-remote-code", "gpu_mem": "0.90", "startup_timeout": "2400",
"run_timeout": "14400",
"model_cache": "/raid/models", "arm_mode": "concurrent", "max_model_len": "auto"},
# GLM-5.2-FP8 (~744GB) needs the whole 8-GPU node, so it runs SEQUENTIAL
# (arm A then arm B; gpu_b unused) unlike the concurrent half-node legs.
# vLLM recipe: TP=8, glm47/glm45, --kv-cache-dtype fp8_e4m3. SMG passes
# glm47_moe explicitly (org-prefixed served name won't auto-detect).
{"name": "glm-5.2", "runner": "blackwell", "tp": 8,
"gpu_a": "0,1,2,3,4,5,6,7", "gpu_b": "", # sequential: arm B reuses GPU_A (whole node)
"model": "zai-org/GLM-5.2-FP8", "bfcl_model": "zai-org/GLM-5.2-FP8-FC",
# GLM-5.3-Flash: 321B/18B-active native-FP8 multimodal MoE (~306GiB, KDA +
# NoPE sparse MLA). Fits a TP=4 half-node, so unlike GLM-5.2-FP8 (~744GB,
# whole node) it runs concurrent like the other Blackwell legs. vLLM recipe
# (vllm-project/recipes models/zai-org/GLM-5.3-Flash.yaml): glm47/glm45,
# --kv-cache-dtype fp8 on Blackwell, FlashInfer >=0.6.18. Support is on
# vLLM main only (vllm-project/vllm#53906, 2026-09-03), so it reuses the
# deepseek-v4.1 leg's per-commit main wheel. SMG passes glm47_moe
# explicitly (org-prefixed served name won't auto-detect).
{"name": "glm-5.3-flash", "runner": "blackwell", "tp": 4,
"gpu_a": "0,1,2,3", "gpu_b": "4,5,6,7",
"model": "zai-org/GLM-5.3-Flash", "bfcl_model": "zai-org/GLM-5.3-Flash-FC",
"vllm_tool": "glm47", "vllm_reason": "glm45",
"smg_tool": "glm47_moe", "smg_reason": "glm45",
# FP8 load + DeepGEMM warmup on the largest leg: generous startup ceiling.
"vllm_extra": "--trust-remote-code --kv-cache-dtype fp8_e4m3", "gpu_mem": "0.90", "startup_timeout": "3000",
# Sequential: both arms share the 360m job, healthy total 2h52m on 2026-08-05.
"vllm_extra": "--trust-remote-code --kv-cache-dtype fp8",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: worth double-checking --kv-cache-dtype fp8 against the sparse-MLA path before the first real run. The deepseek-v4.1 leg 50 lines up had to drop its explicit kv-cache flag precisely because vLLM's sparse-MLA path resolves auto → fp8_ds_mla itself and rejects/ignores a caller-chosen dtype; GLM-5.3-Flash is described here as NoPE sparse MLA too, so plain fp8 may hit the same mismatch on this wheel (the recipe's flag list can lag the model PR that changed the resolution). Cheap to confirm from the arm-A log line that prints the resolved KV dtype on a only=glm-5.3-flash dispatch; if it errors, auto (i.e. omit the flag) is the safe value. Both arms get vllm_extra, so a failure here takes down the whole leg rather than skewing the A/B.

"vllm_commit": "a31ec3a68bbbedd4b5d59490c762bcf32abf1a17",
"vllm_version": "0.29.1rc1.dev185+ga31ec3a68",
# Recipe sets VLLM_ENGINE_READY_TIMEOUT_S=3600: FP8 load + KDA/sparse-MLA
# JIT on a new arch. Kept at the V4.1 leg's ceiling.
"gpu_mem": "0.90", "startup_timeout": "3600",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Important: the recipe's VLLM_ENGINE_READY_TIMEOUT_S=3600 isn't actually applied — nothing in this repo ever sets that variable (grep -rn VLLM_ENGINE_READY_TIMEOUT hits only this comment). startup_timeout flows to BFCL_STARTUP_TIMEOUT, which is only the external health poll in scripts/bfcl/launch_arm.sh (wait_http/wait_grpc, line 119/140). vLLM's engine-core readiness handshake is enforced inside the server process against its own env default (600s), so if GLM-5.3-Flash's FP8 load + KDA/sparse-MLA JIT takes longer than that, the engine aborts itself and the 3600s poll just watches a dead process — the leg fails at startup regardless of this value. The recipe raising the ceiling to 3600 is evidence this model does exceed the default.

Fix is to export it into the server's environment, e.g. in the "Launch arms + run official BFCL A/B" step env: block:

VLLM_ENGINE_READY_TIMEOUT_S: ${{ matrix.startup_timeout }}

(same for nightly-tau2.yml; the launch scripts setsid-detach the servers from this step's shell, so step-level env is inherited by both arms.)

# Per-arm cap carried over from the GLM-5.2 leg (healthy 2h52m for both
# arms serially); arms are concurrent here, so this is generous.
"run_timeout": "7200",
"model_cache": "/raid/models", "arm_mode": "sequential", "max_model_len": "auto"},
"model_cache": "/raid/models", "arm_mode": "concurrent", "max_model_len": "auto"},
]
only = os.environ.get("ONLY", "")
valid_names = [l["name"] for l in legs]
Expand All @@ -244,7 +253,7 @@ jobs:
matrix: ${{ fromJSON(needs.setup.outputs.matrix) }}
runs-on: ${{ matrix.runner }}
# Generous: the nightly set adds multi_turn (state simulation), much slower
# than the single-turn AST/live categories. The sequential glm-5.2 leg needs
# than the single-turn AST/live categories. The sequential deepseek-v4.1 leg needs
# the most headroom — two whole-node arms (startup + scoring) run serially,
# not concurrently like the half-node legs. Ceiling only; fast legs finish early.
timeout-minutes: 360
Expand Down
31 changes: 17 additions & 14 deletions .github/workflows/nightly-tau2.yml
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@
# the ONLY variable is the frontend (tokenization + tool/reasoning parsing).
#
# Runs nightly across retail, airline, and telecom domains over a 6-model matrix
# (qwen3.8 + gpt-oss on H100; deepseek-v4.1, minimax-m2.7, kimi-k2.6, glm-5.2 on
# (qwen3.8 + gpt-oss on H100; deepseek-v4.1, minimax-m2.7, kimi-k2.6, glm-5.3-flash on
# Blackwell). The 4 Blackwell legs are num_tasks-capped to bound GPU time and
# gpt-5.2 spend (tau2 is far heavier per-leg than bfcl: multi-turn + per-turn API
# spend). On PRs that touch this pipeline ALL legs run on a tiny retail / 1-trial /
Expand All @@ -31,7 +31,7 @@ on:
workflow_dispatch:
inputs:
only:
description: "Run only this matrix leg (qwen3.8|gpt-oss|deepseek-v4.1|minimax-m2.7|kimi-k2.6|glm-5.2); empty = all"
description: "Run only this matrix leg (qwen3.8|gpt-oss|deepseek-v4.1|minimax-m2.7|kimi-k2.6|glm-5.3-flash); empty = all"
required: false
default: ""
model:
Expand Down Expand Up @@ -166,7 +166,7 @@ jobs:
# Concurrent TP=4 per arm (0-3 + 4-7); num_tasks-capped.
# DeepSeek-V4.1-Flash is ~765GB on disk (552B backbone + 196B int8 Engram,
# all GPU-resident), so it outgrows a TP=4 half-node and runs SEQUENTIAL
# on the whole node like glm-5.2. SMG: deepseek_v41 tool/reasoning parsers
# on the whole node (the only such leg). SMG: deepseek_v41 tool/reasoning parsers
# + native renderer (#2524/#2525/#2527). vLLM: V4.1 is on main only
# (vllm-project/vllm#56208 + #56228, 2026-09-10), not in the 0.27.1 CI
# pin, so this leg installs a per-commit main wheel (vllm_commit below).
Expand Down Expand Up @@ -201,18 +201,21 @@ jobs:
"vllm_extra": "--trust-remote-code", "gpu_mem": "0.90", "startup_timeout": "2400",
"model_cache": "/raid/models", "arm_mode": "concurrent",
"max_model_len": "auto", "num_tasks": 30, "run_timeout": "5400"},
# GLM-5.2-FP8 (~744GB) needs the whole 8-GPU node, so it runs SEQUENTIAL
# (arm A then arm B; gpu_b unused) unlike the concurrent half-node legs.
{"name": "glm-5.2", "runner": "blackwell", "tp": 8,
"gpu_a": "0,1,2,3,4,5,6,7", "gpu_b": "",
"model": "zai-org/GLM-5.2-FP8",
# GLM-5.3-Flash (~306GiB FP8) fits a TP=4 half-node, so it runs concurrent,
# unlike the whole-node GLM-5.2-FP8 leg it replaces. Flags per the vLLM
# recipe; per-commit main wheel shared with deepseek-v4.1 (see nightly-bfcl.yml).
{"name": "glm-5.3-flash", "runner": "blackwell", "tp": 4,
"gpu_a": "0,1,2,3", "gpu_b": "4,5,6,7",
"model": "zai-org/GLM-5.3-Flash",
"vllm_tool": "glm47", "vllm_reason": "glm45",
"smg_tool": "glm47_moe", "smg_reason": "glm45",
"vllm_extra": "--trust-remote-code --kv-cache-dtype fp8_e4m3",
"gpu_mem": "0.90", "startup_timeout": "3000", "model_cache": "/raid/models",
"arm_mode": "sequential", "max_model_len": "auto", "num_tasks": 30,
# Sequential: both arms share the 360m job, so half the per-domain budget.
"run_timeout": "2700"},
"vllm_extra": "--trust-remote-code --kv-cache-dtype fp8",
"vllm_commit": "a31ec3a68bbbedd4b5d59490c762bcf32abf1a17",
"vllm_version": "0.29.1rc1.dev185+ga31ec3a68",
"gpu_mem": "0.90", "startup_timeout": "3600", "model_cache": "/raid/models",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Important: same as nightly-bfcl.yml — startup_timeout: 3600 only widens TAU2_STARTUP_TIMEOUT (the wait_http/wait_grpc poll in scripts/tau2/launch_arms.sh), it does not set the recipe's VLLM_ENGINE_READY_TIMEOUT_S, which is enforced inside the vLLM process against its own default. Export VLLM_ENGINE_READY_TIMEOUT_S: ${{ matrix.startup_timeout }} in the "Launch arms + run τ²-bench A/B" step env: block so the engine-side ceiling matches.

"arm_mode": "concurrent", "max_model_len": "auto", "num_tasks": 30,
# Concurrent arms: full per-domain budget like the other half-node legs.
"run_timeout": "5400"},
]
only = os.environ.get("ONLY", "")
valid = [l["name"] for l in legs]
Expand All @@ -235,7 +238,7 @@ jobs:
matrix: ${{ fromJSON(needs.setup.outputs.matrix) }}
runs-on: ${{ matrix.runner }}
# Generous ceiling: multi-turn state simulation is slow, and the sequential
# glm-5.2 leg scores two whole-node arms serially. Ceiling only; fast legs
# deepseek-v4.1 leg scores two whole-node arms serially. Ceiling only; fast legs
# finish early. --max-concurrency keeps the full run within budget.
timeout-minutes: 360
permissions:
Expand Down
11 changes: 6 additions & 5 deletions scripts/bfcl/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,6 +67,7 @@ Key env knobs for `launch_arm.sh`: `BFCL_GPU` (CUDA_VISIBLE_DEVICES, e.g. `0,1`)
| DeepSeek-V4.1-Flash (`deepseek-v4.1`) | `blackwell` | 8 (seq) | `deepseek_v41` / `deepseek_v41` (+`--tokenizer-mode deepseek_v41 --trust-remote-code`; per-commit vLLM main wheel) | `deepseek_v41` / `deepseek_v41` |
| MiniMax-M2.7 (`minimax-m2.7`) | `blackwell` | 4 | `minimax_m2` / `minimax_m2` (+`--trust-remote-code`) | `minimax_m2` / `minimax` |
| Kimi-K2.6 int4 (`kimi-k2.6`) | `blackwell` | 4 | `kimi_k2` / `kimi_k2` (+`--trust-remote-code`) | `kimik2` / `kimi_k25`† |
| GLM-5.3-Flash (`glm-5.3-flash`) | `blackwell` | 4 | `glm47` / `glm45` (+`--trust-remote-code --kv-cache-dtype fp8`; per-commit vLLM main wheel) | `glm47_moe` / `glm45` |

> **gpt-oss has no SMG tool-call-parser.** SMG handles gpt-oss through its harmony
> pipeline (`model_gateway/src/routers/grpc/harmony/`), auto-activated by
Expand All @@ -89,8 +90,8 @@ The nightly (`.github/workflows/nightly-bfcl.yml`) runs the A/B as a GitHub Acti
matrix — one leg per model, `fail-fast: false`, each on its own runner:

- `4-gpu-h100` — Qwen3.8-27B and gpt-oss-120b, TP=2 per arm (GPUs 0,1 + 2,3).
- `blackwell` (B200) — MiniMax-M2.7 and Kimi-K2.6 int4, TP=4 per arm (GPUs 0-3 + 4-7);
DeepSeek-V4.1-Flash and GLM-5.2-FP8 need the whole node (TP=8, arms sequential).
- `blackwell` (B200) — MiniMax-M2.7, Kimi-K2.6 int4 and GLM-5.3-Flash, TP=4 per arm
(GPUs 0-3 + 4-7); DeepSeek-V4.1-Flash needs the whole node (TP=8, arms sequential).

All legs use `max_model_len` **32768**: the `multi_turn` categories emit ~18k-token
prompts that 400'd ("decoder prompt longer than the maximum model length") at 16384.
Expand All @@ -101,7 +102,7 @@ Each leg sets `arm_mode`:

- **concurrent** (the half-node legs) — both arms serve at once on opposite GPU halves;
`run_ab.py` scores them **in parallel** (separate servers/GPUs, no contention) and diffs.
- **sequential** (`deepseek-v4.1`, `glm-5.2`) — for a model that needs the whole node (TP=8)
- **sequential** (`deepseek-v4.1`) — for a model that needs the whole node (TP=8)
so the arms can't coexist: `run_ab.py --score-arm` scores arm A alone → tears it
down → scores arm B alone → `--diff-baseline/--diff-candidate` compares the two
saved score files. Flip a leg's `arm_mode` to enable it.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: "Flip a leg's arm_mode to enable it." is a leftover from when sequential was "in reserve, unused" — now that deepseek-v4.1 ships as a sequential leg, this trailing sentence reads as if the mode still needs enabling. Suggest dropping it (or rewording to "set a leg's arm_mode to sequential to use it").

Expand All @@ -118,8 +119,8 @@ is a tiny non-live subset (`simple_python,irrelevance`) for every leg.
A leg whose model the pinned vLLM release (`scripts/ci_install_vllm.sh`) cannot serve
sets `vllm_commit` + `vllm_version` in its matrix entry; the job then swaps in that
per-commit main wheel from `wheels.vllm.ai/<commit>` for **both** arms (the A/B stays
engine-identical). `deepseek-v4.1` uses this until a vLLM release ships V4.1 and the
CI pin moves; drop the two keys then.
engine-identical). `deepseek-v4.1` and `glm-5.3-flash` share one such wheel until a
vLLM release ships both models and the CI pin moves; drop the keys then.

## Gotchas discovered while bringing this up (read before debugging)

Expand Down
4 changes: 2 additions & 2 deletions scripts/tau2/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -87,7 +87,7 @@ where `<DATA_DIR>` is the path you pass to the **required** `--data-dir` flag (t

Mirrors `nightly-bfcl.yml`'s 6-leg matrix. The 2 H100 legs run full task sets; the 4
Blackwell legs are capped at `num_tasks=30`/domain (tau2 is multi-turn + spends
gpt-5.2 per turn). `glm-5.2` runs **sequential** (whole 8-GPU node per arm — `run_ab.py
gpt-5.2 per turn). `deepseek-v4.1` runs **sequential** (whole 8-GPU node per arm — `run_ab.py
--score-arm` each arm, then `--diff`); the rest run both arms concurrently on opposite
GPU halves. On PRs **all** legs run on a tiny retail / 1-trial / few-task subset — a
quick "does each leg launch + parse + score" smoke (the heavy Blackwell legs are
Expand All @@ -100,7 +100,7 @@ dominated by model-load time, serialized by a host lock, so a PR run is not fast
| deepseek-v4.1 | deepseek-ai/DeepSeek-V4.1-Flash | blackwell (8, seq) | `deepseek_v41` / `deepseek_v41` | `deepseek_v41` / `deepseek_v41` |
| minimax-m2.7 | MiniMaxAI/MiniMax-M2.7 | blackwell (4) | `minimax_m2` / `minimax_m2` | `minimax_m2` / `minimax` |
| kimi-k2.6 | moonshotai/Kimi-K2.6 | blackwell (4) | `kimi_k2` / `kimi_k2` | `kimik2` / `kimi_k25` |
| glm-5.2 | zai-org/GLM-5.2-FP8 | blackwell (8, seq) | `glm47` / `glm45` | `glm47_moe` / `glm45` |
| glm-5.3-flash | zai-org/GLM-5.3-Flash | blackwell (4) | `glm47` / `glm45` | `glm47_moe` / `glm45` |

> Dispatch `only=<leg>` runs a single leg; `model=` overrides its weights. SKU ids and
> vLLM parser names may shift; confirm against the installed vLLM build:
Expand Down
Loading