-
Notifications
You must be signed in to change notification settings - Fork 178
ci(bench): move the BFCL and tau2 GLM legs from GLM-5.2-FP8 to GLM-5.3-Flash #2552
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -23,7 +23,7 @@ on: | |
| workflow_dispatch: | ||
| inputs: | ||
| only: | ||
| description: "Run only this matrix leg (qwen3.8|gpt-oss|deepseek-v4.1|minimax-m2.7|kimi-k2.6|glm-5.2); empty = all" | ||
| description: "Run only this matrix leg (qwen3.8|gpt-oss|deepseek-v4.1|minimax-m2.7|kimi-k2.6|glm-5.3-flash); empty = all" | ||
| required: false | ||
| default: "" | ||
| model: | ||
|
|
@@ -162,7 +162,7 @@ jobs: | |
| # Concurrent TP=4 per arm (0-3 + 4-7). | ||
| # DeepSeek-V4.1-Flash is ~765GB on disk (552B backbone + 196B int8 Engram, | ||
| # all GPU-resident), so it outgrows a TP=4 half-node and runs SEQUENTIAL | ||
| # on the whole node like glm-5.2. SMG: deepseek_v41 tool/reasoning parsers | ||
| # on the whole node (the only such leg). SMG: deepseek_v41 tool/reasoning parsers | ||
| # + native renderer (#2524/#2525/#2527). vLLM: V4.1 is on main only | ||
| # (vllm-project/vllm#56208 + #56228, 2026-09-10), not in the 0.27.1 CI | ||
| # pin, so this leg installs a per-commit main wheel (vllm_commit below). | ||
|
|
@@ -208,20 +208,29 @@ jobs: | |
| "vllm_extra": "--trust-remote-code", "gpu_mem": "0.90", "startup_timeout": "2400", | ||
| "run_timeout": "14400", | ||
| "model_cache": "/raid/models", "arm_mode": "concurrent", "max_model_len": "auto"}, | ||
| # GLM-5.2-FP8 (~744GB) needs the whole 8-GPU node, so it runs SEQUENTIAL | ||
| # (arm A then arm B; gpu_b unused) unlike the concurrent half-node legs. | ||
| # vLLM recipe: TP=8, glm47/glm45, --kv-cache-dtype fp8_e4m3. SMG passes | ||
| # glm47_moe explicitly (org-prefixed served name won't auto-detect). | ||
| {"name": "glm-5.2", "runner": "blackwell", "tp": 8, | ||
| "gpu_a": "0,1,2,3,4,5,6,7", "gpu_b": "", # sequential: arm B reuses GPU_A (whole node) | ||
| "model": "zai-org/GLM-5.2-FP8", "bfcl_model": "zai-org/GLM-5.2-FP8-FC", | ||
| # GLM-5.3-Flash: 321B/18B-active native-FP8 multimodal MoE (~306GiB, KDA + | ||
| # NoPE sparse MLA). Fits a TP=4 half-node, so unlike GLM-5.2-FP8 (~744GB, | ||
| # whole node) it runs concurrent like the other Blackwell legs. vLLM recipe | ||
| # (vllm-project/recipes models/zai-org/GLM-5.3-Flash.yaml): glm47/glm45, | ||
| # --kv-cache-dtype fp8 on Blackwell, FlashInfer >=0.6.18. Support is on | ||
| # vLLM main only (vllm-project/vllm#53906, 2026-09-03), so it reuses the | ||
| # deepseek-v4.1 leg's per-commit main wheel. SMG passes glm47_moe | ||
| # explicitly (org-prefixed served name won't auto-detect). | ||
| {"name": "glm-5.3-flash", "runner": "blackwell", "tp": 4, | ||
| "gpu_a": "0,1,2,3", "gpu_b": "4,5,6,7", | ||
| "model": "zai-org/GLM-5.3-Flash", "bfcl_model": "zai-org/GLM-5.3-Flash-FC", | ||
| "vllm_tool": "glm47", "vllm_reason": "glm45", | ||
| "smg_tool": "glm47_moe", "smg_reason": "glm45", | ||
| # FP8 load + DeepGEMM warmup on the largest leg: generous startup ceiling. | ||
| "vllm_extra": "--trust-remote-code --kv-cache-dtype fp8_e4m3", "gpu_mem": "0.90", "startup_timeout": "3000", | ||
| # Sequential: both arms share the 360m job, healthy total 2h52m on 2026-08-05. | ||
| "vllm_extra": "--trust-remote-code --kv-cache-dtype fp8", | ||
| "vllm_commit": "a31ec3a68bbbedd4b5d59490c762bcf32abf1a17", | ||
| "vllm_version": "0.29.1rc1.dev185+ga31ec3a68", | ||
| # Recipe sets VLLM_ENGINE_READY_TIMEOUT_S=3600: FP8 load + KDA/sparse-MLA | ||
| # JIT on a new arch. Kept at the V4.1 leg's ceiling. | ||
| "gpu_mem": "0.90", "startup_timeout": "3600", | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 🔴 Important: the recipe's Fix is to export it into the server's environment, e.g. in the "Launch arms + run official BFCL A/B" step VLLM_ENGINE_READY_TIMEOUT_S: ${{ matrix.startup_timeout }}(same for |
||
| # Per-arm cap carried over from the GLM-5.2 leg (healthy 2h52m for both | ||
| # arms serially); arms are concurrent here, so this is generous. | ||
| "run_timeout": "7200", | ||
| "model_cache": "/raid/models", "arm_mode": "sequential", "max_model_len": "auto"}, | ||
| "model_cache": "/raid/models", "arm_mode": "concurrent", "max_model_len": "auto"}, | ||
| ] | ||
| only = os.environ.get("ONLY", "") | ||
| valid_names = [l["name"] for l in legs] | ||
|
|
@@ -244,7 +253,7 @@ jobs: | |
| matrix: ${{ fromJSON(needs.setup.outputs.matrix) }} | ||
| runs-on: ${{ matrix.runner }} | ||
| # Generous: the nightly set adds multi_turn (state simulation), much slower | ||
| # than the single-turn AST/live categories. The sequential glm-5.2 leg needs | ||
| # than the single-turn AST/live categories. The sequential deepseek-v4.1 leg needs | ||
| # the most headroom — two whole-node arms (startup + scoring) run serially, | ||
| # not concurrently like the half-node legs. Ceiling only; fast legs finish early. | ||
| timeout-minutes: 360 | ||
|
|
||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -8,7 +8,7 @@ | |
| # the ONLY variable is the frontend (tokenization + tool/reasoning parsing). | ||
| # | ||
| # Runs nightly across retail, airline, and telecom domains over a 6-model matrix | ||
| # (qwen3.8 + gpt-oss on H100; deepseek-v4.1, minimax-m2.7, kimi-k2.6, glm-5.2 on | ||
| # (qwen3.8 + gpt-oss on H100; deepseek-v4.1, minimax-m2.7, kimi-k2.6, glm-5.3-flash on | ||
| # Blackwell). The 4 Blackwell legs are num_tasks-capped to bound GPU time and | ||
| # gpt-5.2 spend (tau2 is far heavier per-leg than bfcl: multi-turn + per-turn API | ||
| # spend). On PRs that touch this pipeline ALL legs run on a tiny retail / 1-trial / | ||
|
|
@@ -31,7 +31,7 @@ on: | |
| workflow_dispatch: | ||
| inputs: | ||
| only: | ||
| description: "Run only this matrix leg (qwen3.8|gpt-oss|deepseek-v4.1|minimax-m2.7|kimi-k2.6|glm-5.2); empty = all" | ||
| description: "Run only this matrix leg (qwen3.8|gpt-oss|deepseek-v4.1|minimax-m2.7|kimi-k2.6|glm-5.3-flash); empty = all" | ||
| required: false | ||
| default: "" | ||
| model: | ||
|
|
@@ -166,7 +166,7 @@ jobs: | |
| # Concurrent TP=4 per arm (0-3 + 4-7); num_tasks-capped. | ||
| # DeepSeek-V4.1-Flash is ~765GB on disk (552B backbone + 196B int8 Engram, | ||
| # all GPU-resident), so it outgrows a TP=4 half-node and runs SEQUENTIAL | ||
| # on the whole node like glm-5.2. SMG: deepseek_v41 tool/reasoning parsers | ||
| # on the whole node (the only such leg). SMG: deepseek_v41 tool/reasoning parsers | ||
| # + native renderer (#2524/#2525/#2527). vLLM: V4.1 is on main only | ||
| # (vllm-project/vllm#56208 + #56228, 2026-09-10), not in the 0.27.1 CI | ||
| # pin, so this leg installs a per-commit main wheel (vllm_commit below). | ||
|
|
@@ -201,18 +201,21 @@ jobs: | |
| "vllm_extra": "--trust-remote-code", "gpu_mem": "0.90", "startup_timeout": "2400", | ||
| "model_cache": "/raid/models", "arm_mode": "concurrent", | ||
| "max_model_len": "auto", "num_tasks": 30, "run_timeout": "5400"}, | ||
| # GLM-5.2-FP8 (~744GB) needs the whole 8-GPU node, so it runs SEQUENTIAL | ||
| # (arm A then arm B; gpu_b unused) unlike the concurrent half-node legs. | ||
| {"name": "glm-5.2", "runner": "blackwell", "tp": 8, | ||
| "gpu_a": "0,1,2,3,4,5,6,7", "gpu_b": "", | ||
| "model": "zai-org/GLM-5.2-FP8", | ||
| # GLM-5.3-Flash (~306GiB FP8) fits a TP=4 half-node, so it runs concurrent, | ||
| # unlike the whole-node GLM-5.2-FP8 leg it replaces. Flags per the vLLM | ||
| # recipe; per-commit main wheel shared with deepseek-v4.1 (see nightly-bfcl.yml). | ||
| {"name": "glm-5.3-flash", "runner": "blackwell", "tp": 4, | ||
| "gpu_a": "0,1,2,3", "gpu_b": "4,5,6,7", | ||
| "model": "zai-org/GLM-5.3-Flash", | ||
| "vllm_tool": "glm47", "vllm_reason": "glm45", | ||
| "smg_tool": "glm47_moe", "smg_reason": "glm45", | ||
| "vllm_extra": "--trust-remote-code --kv-cache-dtype fp8_e4m3", | ||
| "gpu_mem": "0.90", "startup_timeout": "3000", "model_cache": "/raid/models", | ||
| "arm_mode": "sequential", "max_model_len": "auto", "num_tasks": 30, | ||
| # Sequential: both arms share the 360m job, so half the per-domain budget. | ||
| "run_timeout": "2700"}, | ||
| "vllm_extra": "--trust-remote-code --kv-cache-dtype fp8", | ||
| "vllm_commit": "a31ec3a68bbbedd4b5d59490c762bcf32abf1a17", | ||
| "vllm_version": "0.29.1rc1.dev185+ga31ec3a68", | ||
| "gpu_mem": "0.90", "startup_timeout": "3600", "model_cache": "/raid/models", | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 🔴 Important: same as |
||
| "arm_mode": "concurrent", "max_model_len": "auto", "num_tasks": 30, | ||
| # Concurrent arms: full per-domain budget like the other half-node legs. | ||
| "run_timeout": "5400"}, | ||
| ] | ||
| only = os.environ.get("ONLY", "") | ||
| valid = [l["name"] for l in legs] | ||
|
|
@@ -235,7 +238,7 @@ jobs: | |
| matrix: ${{ fromJSON(needs.setup.outputs.matrix) }} | ||
| runs-on: ${{ matrix.runner }} | ||
| # Generous ceiling: multi-turn state simulation is slow, and the sequential | ||
| # glm-5.2 leg scores two whole-node arms serially. Ceiling only; fast legs | ||
| # deepseek-v4.1 leg scores two whole-node arms serially. Ceiling only; fast legs | ||
| # finish early. --max-concurrency keeps the full run within budget. | ||
| timeout-minutes: 360 | ||
| permissions: | ||
|
|
||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -67,6 +67,7 @@ Key env knobs for `launch_arm.sh`: `BFCL_GPU` (CUDA_VISIBLE_DEVICES, e.g. `0,1`) | |
| | DeepSeek-V4.1-Flash (`deepseek-v4.1`) | `blackwell` | 8 (seq) | `deepseek_v41` / `deepseek_v41` (+`--tokenizer-mode deepseek_v41 --trust-remote-code`; per-commit vLLM main wheel) | `deepseek_v41` / `deepseek_v41` | | ||
| | MiniMax-M2.7 (`minimax-m2.7`) | `blackwell` | 4 | `minimax_m2` / `minimax_m2` (+`--trust-remote-code`) | `minimax_m2` / `minimax` | | ||
| | Kimi-K2.6 int4 (`kimi-k2.6`) | `blackwell` | 4 | `kimi_k2` / `kimi_k2` (+`--trust-remote-code`) | `kimik2` / `kimi_k25`† | | ||
| | GLM-5.3-Flash (`glm-5.3-flash`) | `blackwell` | 4 | `glm47` / `glm45` (+`--trust-remote-code --kv-cache-dtype fp8`; per-commit vLLM main wheel) | `glm47_moe` / `glm45` | | ||
|
|
||
| > **gpt-oss has no SMG tool-call-parser.** SMG handles gpt-oss through its harmony | ||
| > pipeline (`model_gateway/src/routers/grpc/harmony/`), auto-activated by | ||
|
|
@@ -89,8 +90,8 @@ The nightly (`.github/workflows/nightly-bfcl.yml`) runs the A/B as a GitHub Acti | |
| matrix — one leg per model, `fail-fast: false`, each on its own runner: | ||
|
|
||
| - `4-gpu-h100` — Qwen3.8-27B and gpt-oss-120b, TP=2 per arm (GPUs 0,1 + 2,3). | ||
| - `blackwell` (B200) — MiniMax-M2.7 and Kimi-K2.6 int4, TP=4 per arm (GPUs 0-3 + 4-7); | ||
| DeepSeek-V4.1-Flash and GLM-5.2-FP8 need the whole node (TP=8, arms sequential). | ||
| - `blackwell` (B200) — MiniMax-M2.7, Kimi-K2.6 int4 and GLM-5.3-Flash, TP=4 per arm | ||
| (GPUs 0-3 + 4-7); DeepSeek-V4.1-Flash needs the whole node (TP=8, arms sequential). | ||
|
|
||
| All legs use `max_model_len` **32768**: the `multi_turn` categories emit ~18k-token | ||
| prompts that 400'd ("decoder prompt longer than the maximum model length") at 16384. | ||
|
|
@@ -101,7 +102,7 @@ Each leg sets `arm_mode`: | |
|
|
||
| - **concurrent** (the half-node legs) — both arms serve at once on opposite GPU halves; | ||
| `run_ab.py` scores them **in parallel** (separate servers/GPUs, no contention) and diffs. | ||
| - **sequential** (`deepseek-v4.1`, `glm-5.2`) — for a model that needs the whole node (TP=8) | ||
| - **sequential** (`deepseek-v4.1`) — for a model that needs the whole node (TP=8) | ||
| so the arms can't coexist: `run_ab.py --score-arm` scores arm A alone → tears it | ||
| down → scores arm B alone → `--diff-baseline/--diff-candidate` compares the two | ||
| saved score files. Flip a leg's `arm_mode` to enable it. | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 🟡 Nit: "Flip a leg's |
||
|
|
@@ -118,8 +119,8 @@ is a tiny non-live subset (`simple_python,irrelevance`) for every leg. | |
| A leg whose model the pinned vLLM release (`scripts/ci_install_vllm.sh`) cannot serve | ||
| sets `vllm_commit` + `vllm_version` in its matrix entry; the job then swaps in that | ||
| per-commit main wheel from `wheels.vllm.ai/<commit>` for **both** arms (the A/B stays | ||
| engine-identical). `deepseek-v4.1` uses this until a vLLM release ships V4.1 and the | ||
| CI pin moves; drop the two keys then. | ||
| engine-identical). `deepseek-v4.1` and `glm-5.3-flash` share one such wheel until a | ||
| vLLM release ships both models and the CI pin moves; drop the keys then. | ||
|
|
||
| ## Gotchas discovered while bringing this up (read before debugging) | ||
|
|
||
|
|
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🟡 Nit: worth double-checking
--kv-cache-dtype fp8against the sparse-MLA path before the first real run. Thedeepseek-v4.1leg 50 lines up had to drop its explicit kv-cache flag precisely because vLLM's sparse-MLA path resolvesauto→fp8_ds_mlaitself and rejects/ignores a caller-chosen dtype; GLM-5.3-Flash is described here as NoPE sparse MLA too, so plainfp8may hit the same mismatch on this wheel (the recipe's flag list can lag the model PR that changed the resolution). Cheap to confirm from the arm-A log line that prints the resolved KV dtype on aonly=glm-5.3-flashdispatch; if it errors,auto(i.e. omit the flag) is the safe value. Both arms getvllm_extra, so a failure here takes down the whole leg rather than skewing the A/B.