Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 6 additions & 4 deletions .claude/commands/add-model-hardware.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,8 @@ breakdown. Do **not** invent image tags. Verify them on the registry first.
Don't guess flags or concurrencies. **Deep-research the InferenceX codebase first**, then
the external sources. Read *several* similar files, not just one, and copy what actually runs.

Check `MODELS.md` before choosing a model, scenario, or precision. Do not reintroduce retired coverage; preserve only explicitly documented exceptions. Use active siblings, not files under `deprecated/`.

**A. In-codebase research (primary because this repo is the source of truth):**
```bash
# similar benchmark scripts: same model on other SKUs, AND same SKU on other models
Expand All @@ -50,10 +52,10 @@ grep -nE "run_benchmark_serving|setup_eval_context|wait_for_server_ready|start_g
- **Read multiple sibling scripts** end-to-end for the exact env vars and serve shape (`VLLM_*`,
`SGLANG_*`, device mapping, download/cache handling, `--enforce-eager` vs graph capture,
KV-cache dtype, attention/MoE backend, parsers). These are the truth for each runner.
- **Compare several master-config search spaces** (e.g. `dsv4`, `glm5`, the same model on a
- **Compare several master-config search spaces** (e.g. `dsr1`, `qwen3.5`, the same model on a
sibling SKU) to choose `{tp, ep, dp-attn} × concurrency` combos that fit *this* hardware's
memory. Small-memory SKUs like h100/mi300x go TP8-only, while bigger SKUs add tp4/tp2/DEP.
- **Internalize the fixed-seq-len nuances from the existing configs**: `8k1k`/`1k8k` do **not**
- **Internalize the fixed-seq-len nuances from the existing configs**: `8k1k` runs do **not**
need the full `MAX_MODEL_LEN` (the matrix supplies `isl + osl + slack`), and graph-capture
batch sizes are scaled to concurrency/scenario (and spec-token count for MTP), not maxed.
Copy how sibling scripts/configs already do it.
Expand Down Expand Up @@ -107,7 +109,7 @@ the model's `recipes.vllm.ai` page:
Capture up to the next power of two ≥ `CONC` (≥ `CONC * (1 + NUM_SPEC_TOKENS)` with spec
decoding), capped at vLLM's 2048.
- **`MAX_MODEL_LEN`** is the matrix-supplied scenario value (`isl + osl + slack`). Never
hardcode the full context for 8k1k / 1k8k.
hardcode the full context for 8k1k.
- **Memory headroom.** Bigger checkpoints constrain TP/EP. If the sibling on a smaller-memory
SKU is TP8-only (e.g. h100), match that.

Expand All @@ -117,7 +119,7 @@ Validate as you go: `bash -n <script>`.

Append `<model>-<precision>-<sku>[-<engine>][-mtp]` after the sibling, with the correct
`image`, `model`, `model-prefix`, `runner`, `precision`, `framework`. The **search space** is
`{tp, ep, dp-attn} × concurrency` per scenario (1k1k, 8k1k):
`{tp, ep, dp-attn} × concurrency` per supported scenario from `MODELS.md` (8k1k or AgentX as applicable; 1k1k is only retained for GLM-5.1 B200 TileRT):
- Mirror a sibling's parallelism layouts. Trim concurrency ranges to what the SKU's memory
supports (small-mem SKUs → TP8-only, drop tp2/tp4 and DEP).
- Latency (TP-only) rows should start at conc 1. TEP/DEP rows start higher (they only pay off
Expand Down
6 changes: 3 additions & 3 deletions .claude/commands/nuke.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,8 +11,8 @@ Arguments (`$ARGUMENTS`): `<engine> <target-tag> [filter]`
- `target-tag`. For example, use `v0.22.0` for NVIDIA/CUDA. For SGLang, the NVIDIA and AMD tag
strings usually differ (CUDA `…-cu130` vs ROCm `…-rocm720-mi35x-…`), so confirm
the exact tag per image repo with the user before editing.
- `filter` (optional). Restrict the scope to a model and/or SKU substring (e.g. `kimik2.5`,
`b300`, `minimaxm2.5 mi355x`). If omitted, all matching recipes are in scope.
- `filter` (optional). Restrict the scope to a model and/or SKU substring (e.g. `qwen3.5`,
`b300`, `dsr1 mi355x`). If omitted, all matching recipes are in scope.

## Image repos by engine + vendor

Expand All @@ -24,7 +24,7 @@ Arguments (`$ARGUMENTS`): `<engine> <target-tag> [filter]`
## Grouping rules (NON-NEGOTIABLE)

1. **One PR per `model + precision + SKU` recipe family.** The config-key shape is
`<model>-<precision>-<sku>-<engine>` (e.g. `kimik2.5-int4-b300-vllm`).
`<model>-<precision>-<sku>-<engine>` (e.g. `dsr1-fp8-b300-vllm`).
2. **Fold the `-mtp` (and non-mtp) sibling into the SAME PR** as its base recipe.
This is the *only* thing you may combine.
3. **Never** put two different models, two different precisions, or two different
Expand Down
2 changes: 1 addition & 1 deletion .github/codeowner-signoff-verify-prompt.md
Original file line number Diff line number Diff line change
Expand Up @@ -343,7 +343,7 @@ Verify BOTH:
Name the config/script and line.
- (b) AL VALUE MATCHES THE GOLDEN CURVE. Read the committed golden AL YAML for the
model in `golden_al_distribution/` on the default-branch checkout. Examples include
`qwen3.5_mtp.yaml` and `kimik2.5_eagle3.yaml`. Confirm the pinned AL equals the golden value for that
`qwen3.5_mtp.yaml` and `minimaxm3_eagle3.yaml`. Confirm the pinned AL equals the golden value for that
model, thinking mode, and the config's `num_speculative_tokens` / MTP level (e.g.
qwen3.5 thinking_on with 3 speculative tokens -> 3.39). For TRT-LLM configs, compare
the pinned `TLLM_SPEC_DECODE_FORCE_NUM_ACCEPTED_TOKENS` value PLUS 1 against the
Expand Down
8 changes: 4 additions & 4 deletions .github/workflows/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,9 +55,9 @@ By default, throughput runs for every generated config and eval-only jobs run fo
full-sweep --config-files configs/nvidia-master.yaml
```

**Test all single-node gptoss configurations on B200 with 1k1k sequence lengths:**
**Test all single-node dsr1 configurations on B200 with 8k1k sequence lengths:**
```
full-sweep --single-node --model-prefix gptoss --runner-type b200 --seq-lens 1k1k --config-files configs/nvidia-master.yaml
full-sweep --single-node --model-prefix dsr1 --runner-type b200 --seq-lens 8k1k --config-files configs/nvidia-master.yaml
```

**Test all single-node fp8 precision configs for 8k1k workloads:**
Expand All @@ -72,7 +72,7 @@ full-sweep --single-node --framework trt --runner-type h200 b200-trt --config-fi

**Test specific single-node model on specific hardware with specific sequence lengths:**
```
full-sweep --single-node --model-prefix dsr1 --runner-type b200 --precision fp4 --framework sglang --seq-lens 1k1k 8k1k --config-files configs/nvidia-master.yaml
full-sweep --single-node --model-prefix dsr1 --runner-type b200 --precision fp4 --framework sglang --seq-lens 8k1k --config-files configs/nvidia-master.yaml
```

**Limit concurrency and parallelism for faster testing:**
Expand Down Expand Up @@ -134,7 +134,7 @@ test-config --config-keys dsr1* --config-files configs/nvidia-master.yaml

**Mix exact keys and patterns:**
```
test-config --config-keys dsr1-fp4-b200-sglang gptoss* --config-files configs/nvidia-master.yaml
test-config --config-keys dsr1-fp4-b200-sglang qwen3.5* --config-files configs/nvidia-master.yaml
```

**Override concurrency for targeted testing:**
Expand Down
42 changes: 15 additions & 27 deletions .github/workflows/claude.yml
Original file line number Diff line number Diff line change
Expand Up @@ -131,7 +131,7 @@ jobs:

**Specify concurrency and sequence length:**
```
generate-cli-command: "full-sweep --config-files configs/nvidia-master.yaml --single-node --model-prefix dsr1 --min-conc 4 --max-conc 4 --seq-lens 1k1k"
generate-cli-command: "full-sweep --config-files configs/nvidia-master.yaml --single-node --model-prefix dsr1 --min-conc 4 --max-conc 4 --seq-lens 8k1k"
```

**Test specific config keys (MUST USE `--conc`):**
Expand All @@ -142,7 +142,7 @@ jobs:
**IMPORTANT: Keep runs precise and efficient:**
- Use `full-sweep` with filter flags to narrow down the benchmark scope - "full-sweep" does NOT mean running everything
- When using `full-sweep`, you must use `--min-conc` and `--max-conc` together to specify a single concurrency value. Unless prompted otherwise, use `--min-conc 4 --max-conc 4`
- When using `full-sweep`, you can use `--seq-lens` to specify sequence lengths (choices: 1k1k, 8k1k). Unless prompted otherwise, use `--seq-lens 1k1k`
- For fixed-sequence runs, use `--seq-lens 8k1k`. Use `1k1k` only for the explicitly retained GLM-5.1 B200 TileRT config. Consult `MODELS.md` for model-specific scenario eligibility.
- Use `test-config` ONLY when given specific config keys to test - Use `--config-files`, `--config-keys`, and `--conc` flags ONLY
- Always filter by specific models, frameworks, precision, conc, or config keys when possible

Expand Down Expand Up @@ -193,7 +193,7 @@ jobs:

**How to map a natural-language request to inputs:**
The user will say something like "profile sglang b200 deepseek fp4 conc=4". Parse it as:
- Model: "deepseek" / "dsr1" → model-prefix `dsr1`; "gptoss" → `gptoss`; "qwen" → `qwen3.5`
- Model: resolve the requested model against the current active master configs and `MODELS.md`; for example, `dsr1` means DeepSeek-R1 and `qwen3.5` means Qwen3.5. Reject retired model/scenario requests and preserve documented exceptions.
- Precision: "fp4" / "fp8" / "bf16"
- Runner/hardware: "b200", "h200", "h100", "mi300x", "mi325x", "mi355x", etc.
- Framework: must be "sglang" (reject if not)
Expand All @@ -203,8 +203,8 @@ jobs:
Choose config-file: NVIDIA runners (b200, h200, h100, gb200, gb300) → `nvidia-master.yaml`; AMD runners (mi300x, mi325x, mi355x) → `amd-master.yaml`

**Available SGLang config keys:**
NVIDIA: `dsr1-fp4-b200-sglang`, `dsr1-fp8-b200-sglang`, `dsr1-fp8-h200-sglang`, `qwen3.5-bf16-b200-sglang`
AMD: `dsr1-fp4-mi355x-sglang`, `dsr1-fp8-mi300x-sglang`, `dsr1-fp8-mi325x-sglang`, `dsr1-fp8-mi355x-sglang`, `qwen3.5-bf16-mi355x-sglang`, `qwen3.5-fp8-mi355x-sglang`
NVIDIA: `dsr1-fp4-b200-sglang`, `dsr1-fp8-b200-sglang`, `dsr1-fp8-h200-sglang`, `qwen3.5-fp8-b200-sglang`
AMD: `dsr1-fp4-mi355x-sglang`, `dsr1-fp8-mi300x-sglang`, `dsr1-fp8-mi325x-sglang`, `dsr1-fp8-mi355x-sglang`, `qwen3.5-fp8-mi355x-sglang`

**Examples:**
- "profile sglang b200 deepseek fp4 conc=4" → `config-key: dsr1-fp4-b200-sglang`, `config-file: configs/nvidia-master.yaml`, `conc: 4`
Expand Down Expand Up @@ -585,27 +585,15 @@ jobs:
## Model Prefix Validation:
When reviewing changes to `configs/*-master.yaml` files, verify that ALL config keys use valid model prefixes.

**Valid model prefixes:**
- `dsr1` - DeepSeek R1 models
- `gptoss` - GPT-OSS models
Use `MODELS.md` and the active master configs as the source of truth for supported models, scenarios, precisions, and documented exceptions. Do not use a hard-coded historical prefix allowlist.

**Invalid model prefixes (will break frontend):**
- `dsr1-fp8` - INVALID: precision should NOT be part of the model prefix
- `dsr1-fp4` - INVALID: precision should NOT be part of the model prefix
- Any other prefix not in the valid list above
**Config identity:**
- `model-prefix` identifies the model family and does not include precision (for example, `dsr1`, not `dsr1-fp8`).
- Config keys begin with `{model-prefix}-{precision}-{hardware}-{framework}` and may have scenario or recipe suffixes.
- Example: `dsr1-fp8-gb200-vllm` has model-prefix `dsr1` and precision `fp8`.

**Config key format:**
Config keys follow the pattern: `{model-prefix}-{precision}-{hardware}-{framework}`
Example valid keys:
- `dsr1-fp8-gb200-vllm` (model-prefix=dsr1, precision=fp8)
- `gptoss-fp4-h200-sglang` (model-prefix=gptoss, precision=fp4)

**Why this matters:**
The frontend expects specific model prefixes (`dsr1` or `gptoss`) to display benchmark results correctly. Using invalid prefixes like `dsr1-fp8` will cause the frontend to fail to display results, breaking the user experience.

**Validation:**
When reviewing config additions or changes, check that the first segment of config keys (before the first `-`) is either `dsr1` or `gptoss`.

If a config key uses an invalid model prefix:
- This is a 🔴 **BLOCKING** issue
- Comment: "Invalid model prefix detected. The frontend only supports `dsr1` and `gptoss` as model prefixes. Using other prefixes like `dsr1-fp8` will break the frontend and prevent benchmark results from being displayed. Please use the format `{valid-prefix}-{precision}-{hardware}-{framework}` where valid-prefix is either `dsr1` or `gptoss`."
**Deprecation validation:**
- Reject new active entries for retired models, scenarios, or precisions unless `MODELS.md` explicitly documents an exception.
- Preserve the GLM-5.1 B200 TileRT exception and the conditional non-speculative AgentX policy; do not infer retirement from the existence of a speculative counterpart.
- Deprecated entries belong only in `configs/deprecated/amd-master.yaml` or `configs/deprecated/nvidia-master.yaml`, following `AGENTS.md`. Historical entries in these archives are not active submissions.
- For a newly supported model, require matching updates to `MODELS.md` and `MODELS_zh.md` rather than rejecting it against an obsolete model list.
2 changes: 1 addition & 1 deletion .github/workflows/run-sweep.yml
Original file line number Diff line number Diff line change
Expand Up @@ -296,7 +296,7 @@ jobs:
- fp4: FP4 precision
- mtp, eagle, eagle3: speculative decoding method
- sglang, vllm, dynamo-vllm: runtime framework
- model criterion: matching configured model family
- model criterion: matching configured model family; legacy criteria remain in the shared priority taxonomy for retained SPEED-Bench collectors and do not authorize deprecated master coverage
- checklist-complete: PR checklist is satisfied
- patchwork: modified upstream engine or runtime source

Expand Down
9 changes: 9 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,6 +50,15 @@ Then validate it in the receiving script after sourcing the shared helper:
check_env_vars IS_MULTINODE MODEL_NAME PRECISION
```

## Deprecating benchmark configs

- Move deprecated entries out of the active master config into [`configs/deprecated/amd-master.yaml`](configs/deprecated/amd-master.yaml) or [`configs/deprecated/nvidia-master.yaml`](configs/deprecated/nvidia-master.yaml), matching the vendor. These are the only deprecated master-config files; do not create separate files per model, scenario, or deprecation.
- For a partial deprecation, archive only the retired scenarios and retain the supported scenarios in the active entry. Preserve archived settings and explanatory comments; do not update historical image pins or runners during archival.
- Keep every archive key unique. If a key already exists with different settings, preserve both versions with a descriptive suffix on the historical key and a comment recording its original config key. Existing colliding 1k1k versions use `-deprecated-1k1k`. Never overwrite an archived version or add duplicate YAML keys.
- Check retirement statements in [`MODELS.md`](MODELS.md) against active configs and script routing in the same PR, and update `MODELS.md` plus `MODELS_zh.md` together. Preserve explicitly documented exceptions and conditional retirement policies; do not treat planned retirement as completed.
- Remove unused retired-model branches from launchers and runtime settings, and update workflow/agent guidance that still recommends retired coverage. Audit callers before removing shared helpers; retained SPEED-Bench collectors and historical result readers may still need model-specific support.
- Keep these archives out of active sweep inputs. Follow the existing benchmark-script archival convention, moving retired scripts into the sibling `deprecated/` directory only when no active config still uses them.

## Runner launchers (one file per pool)

- The reusable workflows run `bash ./runners/launch_${RUNNER_NAME%%_*}.sh`. The runner-name prefix before the first underscore is the only routing key, so each self-hosted pool maps to exactly one `runners/launch_<pool>.sh`, and every `runners/launch_*.sh` must be the launcher of a pool listed in [`configs/runners.yaml`](configs/runners.yaml). For example, runner `b200-nscale-slurm_03` runs `runners/launch_b200-nscale-slurm.sh`. See [Stage 4 in `docs/architecture.md`](docs/architecture.md#stage-4-launcher-and-runtime-execution).
Expand Down
Loading
Loading