docs: complete sglang→vllm rename across docs/ + README - #38
Conversation
Split from PR #18 (gcl/clean-sglang). One of 4 PRs splitting the original PR #18 by content area (docs / examples / tests+CI / core runtime). See pr18-split-plan.md. 53 files. Pure documentation: no runtime code change. What this PR does: - docs/en/ + docs/zh/ (49 files): rename sglang-config.md → vllm-config.md (1 each side); update all --sglang-* CLI flag examples to their --vllm-* equivalents; rename references to SGLang as the rollout backend → vLLM; fix three broken upstream URLs (each replacement verified to load): * blog.vllm.ai/2025/10/26/distributed-batch-invariance.html → vllm.ai/blog/2025-11-10-bitwise-consistent-train-inference * docs.vllm.ai/.../features/spec_decode.html (301 → 404) → docs.vllm.ai/.../features/speculative_decoding/ * github.com/vllm-project/vllm-router (404; repo does not exist) → github.com/vllm-project/router (the actual router repo) - docs/_static/image/ (3 files): rename sglang_config.png → vllm_config.png; refresh arch.png to the vllm-router architecture. - docs/en/advanced/reproducibility.md and docs/zh/advanced/reproducibility.md: the en and zh files had their content swapped (en had Chinese body, zh had English body). Restored the correct mapping. - docs/{en,zh}/blogs/introducing_slime.md: remove the "Special thanks to the AMD GenAI - Foundation Model Team for Day-1 AMD hardware support" line. With AMD support removed in #36, the acknowledgement implied ongoing support that no longer exists. - README.md / README_zh.md: expanded vLLM router argument prefix explanation (--router-* for vllm-router native flags vs --vllm-router-* for vime-side orchestration knobs); remove stale link to docs/zh/advanced/ slime_vllm_backend_design_v1.md which was deleted in commit fe700db. This PR is independent of the other splits and can land at any time without coordination. See pr18-split-plan.md §"合并顺序无强约束". Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Code Review
This pull request comprehensively updates the English and Chinese documentation to transition the framework from SGLang-native to vLLM-native, replacing SGLang references, CLI flags, and architecture details with their vLLM and vllm-router equivalents. The reviewer identified several instances in the documentation and code examples where router configuration flags (such as --router-ip, --router-port, and --router-request-timeout-secs) and argument attributes (such as args.router_ip and args.router_port) were incorrectly prefixed with vllm- or vllm_. Actionable suggestions were provided to use the correct prefix-free versions as defined in the codebase.
|
|
||
| 1. **Megatron arguments**: Vime reads all arguments in Megatron. You can configure Megatron by passing arguments like `--tensor-model-parallel-size 2`. | ||
| 2. **vLLM arguments**: vLLM server and engine options are exposed with a `--vllm-` prefix (for example, `--vllm-gpu-memory-utilization`). Router-related options use the `--router-` prefix. See [slime/backends/vllm_utils/arguments.py](slime/backends/vllm_utils/arguments.py) for the full surface. | ||
| 2. **vLLM arguments**: vLLM server and engine options are exposed with a `--vllm-` prefix (for example, `--vllm-gpu-memory-utilization`). Router options live under two prefixes: vllm-router's native options are passed with `--router-` (for example, `--router-policy round_robin`), while Vime-side orchestration knobs that tell Vime *where* the router lives use `--vllm-router-` (`--vllm-router-ip`, `--vllm-router-port`, `--vllm-router-request-timeout-secs`). See [slime/backends/vllm_utils/arguments.py](slime/backends/vllm_utils/arguments.py) for the full surface. |
There was a problem hiding this comment.
The router configuration flags are --router-ip, --router-port, and --router-request-timeout-secs (without the --vllm- prefix), as defined in slime/backends/vllm_utils/arguments.py. Please update the documentation to reflect the correct flags.
| 2. **vLLM arguments**: vLLM server and engine options are exposed with a `--vllm-` prefix (for example, `--vllm-gpu-memory-utilization`). Router options live under two prefixes: vllm-router's native options are passed with `--router-` (for example, `--router-policy round_robin`), while Vime-side orchestration knobs that tell Vime *where* the router lives use `--vllm-router-` (`--vllm-router-ip`, `--vllm-router-port`, `--vllm-router-request-timeout-secs`). See [slime/backends/vllm_utils/arguments.py](slime/backends/vllm_utils/arguments.py) for the full surface. | |
| 2. **vLLM arguments**: vLLM server and engine options are exposed with a `--vllm-` prefix (for example, `--vllm-gpu-memory-utilization`). Router options live under the `--router-` prefix: vllm-router's native options are passed with `--router-` (for example, `--router-policy round_robin`), and Vime-side orchestration knobs that tell Vime *where* the router lives also use `--router-` (`--router-ip`, `--router-port`, `--router-request-timeout-secs`). See [slime/backends/vllm_utils/arguments.py](slime/backends/vllm_utils/arguments.py) for the full surface. |
|
|
||
| 1. **Megatron 参数**:Vime 会读取 Megatron 中的全部参数,可通过传入如 `--tensor-model-parallel-size 2` 的方式配置 Megatron; | ||
| 2. **vLLM 参数**:vLLM server 与 engine 相关选项以 `--vllm-` 为前缀(例如 `--vllm-gpu-memory-utilization`);路由相关选项使用 `--router-` 前缀。完整参数见 [slime/backends/vllm_utils/arguments.py](slime/backends/vllm_utils/arguments.py)。 | ||
| 2. **vLLM 参数**:vLLM server 与 engine 相关选项以 `--vllm-` 为前缀(例如 `--vllm-gpu-memory-utilization`)。路由相关选项分两类前缀:vllm-router 自身的选项以 `--router-` 传入(例如 `--router-policy round_robin`),Vime 侧用于告诉 Vime *router 在哪里* 的编排参数则以 `--vllm-router-` 为前缀(`--vllm-router-ip`、`--vllm-router-port`、`--vllm-router-request-timeout-secs`)。完整参数见 [slime/backends/vllm_utils/arguments.py](slime/backends/vllm_utils/arguments.py)。 |
There was a problem hiding this comment.
The router configuration flags are --router-ip, --router-port, and --router-request-timeout-secs (without the --vllm- prefix), as defined in slime/backends/vllm_utils/arguments.py. Please update the documentation to reflect the correct flags.
| 2. **vLLM 参数**:vLLM server 与 engine 相关选项以 `--vllm-` 为前缀(例如 `--vllm-gpu-memory-utilization`)。路由相关选项分两类前缀:vllm-router 自身的选项以 `--router-` 传入(例如 `--router-policy round_robin`),Vime 侧用于告诉 Vime *router 在哪里* 的编排参数则以 `--vllm-router-` 为前缀(`--vllm-router-ip`、`--vllm-router-port`、`--vllm-router-request-timeout-secs`)。完整参数见 [slime/backends/vllm_utils/arguments.py](slime/backends/vllm_utils/arguments.py)。 | |
| 2. **vLLM 参数**:vLLM server 与 engine 相关选项以 `--vllm-` 为前缀(例如 `--vllm-gpu-memory-utilization`)。路由相关选项统一使用 `--router-` 前缀:vllm-router 自身的选项以 `--router-` 传入(例如 `--router-policy round_robin`),Vime 侧用于告诉 Vime *router 在哪里* 的编排参数也以 `--router-` 为前缀(`--router-ip`、`--router-port`、`--router-request-timeout-secs`)。完整参数见 [slime/backends/vllm_utils/arguments.py](slime/backends/vllm_utils/arguments.py)。 |
| output = await post( | ||
| f"http://{args.sglang_router_ip}:{args.sglang_router_port}/generate", | ||
| f"http://{args.vllm_router_ip}:{args.vllm_router_port}/inference/v1/generate", | ||
| { | ||
| "text": sample.prompt, | ||
| "sampling_params": sampling_params, | ||
| "model": args.hf_checkpoint, | ||
| "token_ids": prompt_token_ids, | ||
| "sampling_params": {"max_tokens": sampling_params["max_new_tokens"]}, | ||
| } | ||
| ) |
There was a problem hiding this comment.
The router configuration attributes on args are args.router_ip and args.router_port (without the vllm_ prefix), as defined in slime/backends/vllm_utils/arguments.py and used in slime/rollout/vllm_rollout.py. Please update the code example to use the correct attributes.
| output = await post( | |
| f"http://{args.sglang_router_ip}:{args.sglang_router_port}/generate", | |
| f"http://{args.vllm_router_ip}:{args.vllm_router_port}/inference/v1/generate", | |
| { | |
| "text": sample.prompt, | |
| "sampling_params": sampling_params, | |
| "model": args.hf_checkpoint, | |
| "token_ids": prompt_token_ids, | |
| "sampling_params": {"max_tokens": sampling_params["max_new_tokens"]}, | |
| } | |
| ) | |
| output = await post( | |
| f"http://{args.router_ip}:{args.router_port}/inference/v1/generate", | |
| { | |
| "model": args.hf_checkpoint, | |
| "token_ids": prompt_token_ids, | |
| "sampling_params": {"max_tokens": sampling_params["max_new_tokens"]}, | |
| } | |
| ) |
| slime uses [vllm-router](https://github.com/vllm-project/router) to manage the vLLM engines during the training process. You can configure the address of the router using `--vllm-router-ip` and `--vllm-router-port`. If not configured, a router will be started by default within the cluster. | ||
|
|
||
| After starting, all SGLang servers will register with the router via the `/add_worker` endpoint. When actually generating data, you only need to send HTTP requests to the router, which will perform load balancing and forward the requests to the servers. | ||
| After starting, all vLLM engines will register with the router. When actually generating data, you only need to send HTTP requests to the router, which will perform load balancing and forward the requests to the engines. | ||
|
|
||
| When you configure an external router using `--sglang-router-ip` and `--sglang-router-port`, slime will not start an internal router. Instead, it will register all its servers with this external router. You can then use this external router's address to implement more complex data generation workflows. Note that the router supports OpenAI-compatible APIs. | ||
| When you configure an external router using `--vllm-router-ip` and `--vllm-router-port`, slime will not start an internal router. Instead, it will register all its engines with this external router. You can then use this external router's address to implement more complex data generation workflows. Note that the router supports OpenAI-compatible APIs. |
There was a problem hiding this comment.
The router configuration flags are --router-ip and --router-port (without the --vllm- prefix), as defined in slime/backends/vllm_utils/arguments.py. Please update the documentation to reflect the correct flags.
| slime uses [vllm-router](https://github.com/vllm-project/router) to manage the vLLM engines during the training process. You can configure the address of the router using `--vllm-router-ip` and `--vllm-router-port`. If not configured, a router will be started by default within the cluster. | |
| After starting, all SGLang servers will register with the router via the `/add_worker` endpoint. When actually generating data, you only need to send HTTP requests to the router, which will perform load balancing and forward the requests to the servers. | |
| After starting, all vLLM engines will register with the router. When actually generating data, you only need to send HTTP requests to the router, which will perform load balancing and forward the requests to the engines. | |
| When you configure an external router using `--sglang-router-ip` and `--sglang-router-port`, slime will not start an internal router. Instead, it will register all its servers with this external router. You can then use this external router's address to implement more complex data generation workflows. Note that the router supports OpenAI-compatible APIs. | |
| When you configure an external router using `--vllm-router-ip` and `--vllm-router-port`, slime will not start an internal router. Instead, it will register all its engines with this external router. You can then use this external router's address to implement more complex data generation workflows. Note that the router supports OpenAI-compatible APIs. | |
| slime uses [vllm-router](https://github.com/vllm-project/router) to manage the vLLM engines during the training process. You can configure the address of the router using `--router-ip` and `--router-port`. If not configured, a router will be started by default within the cluster. | |
| After starting, all vLLM engines will register with the router. When actually generating data, you only need to send HTTP requests to the router, which will perform load balancing and forward the requests to the engines. | |
| When you configure an external router using `--router-ip` and `--router-port`, slime will not start an internal router. Instead, it will register all its engines with this external router. You can then use this external router's address to implement more complex data generation workflows. Note that the router supports OpenAI-compatible APIs. |
| prompt_token_ids = TOKENIZER(sample.prompt, add_special_tokens=False)["input_ids"] | ||
| output = await post( | ||
| f"http://{args.sglang_router_ip}:{args.sglang_router_port}/generate", | ||
| f"http://{args.vllm_router_ip}:{args.vllm_router_port}/inference/v1/generate", | ||
| { | ||
| "text": sample.prompt, | ||
| "sampling_params": sampling_params, | ||
| "model": args.hf_checkpoint, | ||
| "token_ids": prompt_token_ids, | ||
| "sampling_params": {"max_tokens": sampling_params["max_new_tokens"]}, | ||
| } |
There was a problem hiding this comment.
The router configuration attributes on args are args.router_ip and args.router_port (without the vllm_ prefix), as defined in slime/backends/vllm_utils/arguments.py and used in slime/rollout/vllm_rollout.py. Please update the code example to use the correct attributes.
| prompt_token_ids = TOKENIZER(sample.prompt, add_special_tokens=False)["input_ids"] | |
| output = await post( | |
| f"http://{args.sglang_router_ip}:{args.sglang_router_port}/generate", | |
| f"http://{args.vllm_router_ip}:{args.vllm_router_port}/inference/v1/generate", | |
| { | |
| "text": sample.prompt, | |
| "sampling_params": sampling_params, | |
| "model": args.hf_checkpoint, | |
| "token_ids": prompt_token_ids, | |
| "sampling_params": {"max_tokens": sampling_params["max_new_tokens"]}, | |
| } | |
| output = await post( | |
| f"http://{args.router_ip}:{args.router_port}/inference/v1/generate", | |
| { | |
| "model": args.hf_checkpoint, | |
| "token_ids": prompt_token_ids, | |
| "sampling_params": {"max_tokens": sampling_params["max_new_tokens"]}, | |
| } | |
| ) |
|
|
||
| slime 会用 [sglang-router](https://github.com/sgl-project/sglang/tree/main/sgl-router) 来管理训练过程中的 sglang server。可以通过 `--sglang-router-ip` 与 `--sglang-router-port` 来配置 [sglang-router](https://github.com/sgl-project/sglang/tree/main/sgl-router) 的地址。如果不进行配置,则会在集群中默认启动一个 router。 | ||
| slime 会用 [vllm-router](https://github.com/vllm-project/router) 来管理训练过程中的 vLLM 引擎。可以通过 `--vllm-router-ip` 与 `--vllm-router-port` 来配置 router 的地址。如果不进行配置,则会在集群中默认启动一个 router。 | ||
|
|
||
| 所有的 sglang server 在启动后,会通过 `/add_worker` 申请加入 router。在实际进行数据生成的时候,只需要向 router 发送 http 请求,router 会进行 load balancing 操作,将请求转发给 server 们。 | ||
| 所有的 vLLM 引擎在启动后会注册到 router。在实际进行数据生成的时候,只需要向 router 发送 http 请求,router 会进行 load balancing 操作,将请求转发给引擎。 | ||
|
|
There was a problem hiding this comment.
The router configuration flags are --router-ip and --router-port (without the --vllm- prefix), as defined in slime/backends/vllm_utils/arguments.py. Please update the documentation to reflect the correct flags.
| slime 会用 [sglang-router](https://github.com/sgl-project/sglang/tree/main/sgl-router) 来管理训练过程中的 sglang server。可以通过 `--sglang-router-ip` 与 `--sglang-router-port` 来配置 [sglang-router](https://github.com/sgl-project/sglang/tree/main/sgl-router) 的地址。如果不进行配置,则会在集群中默认启动一个 router。 | |
| slime 会用 [vllm-router](https://github.com/vllm-project/router) 来管理训练过程中的 vLLM 引擎。可以通过 `--vllm-router-ip` 与 `--vllm-router-port` 来配置 router 的地址。如果不进行配置,则会在集群中默认启动一个 router。 | |
| 所有的 sglang server 在启动后,会通过 `/add_worker` 申请加入 router。在实际进行数据生成的时候,只需要向 router 发送 http 请求,router 会进行 load balancing 操作,将请求转发给 server 们。 | |
| 所有的 vLLM 引擎在启动后会注册到 router。在实际进行数据生成的时候,只需要向 router 发送 http 请求,router 会进行 load balancing 操作,将请求转发给引擎。 | |
| slime 会用 [vllm-router](https://github.com/vllm-project/router) 来管理训练过程中的 vLLM 引擎。可以通过 `--router-ip` 与 `--router-port` 来配置 router 的地址。如果不进行配置,则会在集群中默认启动一个 router。 | |
| 所有的 vLLM 引擎在启动后会注册到 router。在实际进行数据生成的时候,只需要向 router 发送 http 请求,router 会进行 load balancing 操作,将请求转发给引擎。 | |
| 当通过 `--router-ip` 与 `--router-port` 来配置传入一个外部的 router,此时 slime 不再会在内部启动一个 router,而是会把所有的引擎都注册在这个外部 router 上。这时可以利用这个外部的 router 地址来实现更复杂的数据生成流程。注意 router 是支持 openai compatible api 的。 |
There was a problem hiding this comment.
The router configuration flags are
--router-ipand--router-port(without the--vllm-prefix), as defined inslime/backends/vllm_utils/arguments.py. Please update the documentation to reflect the correct flags.Suggested change
slime 会用 sglang-router 来管理训练过程中的 sglang server。可以通过--sglang-router-ip与--sglang-router-port来配置 sglang-router 的地址。如果不进行配置,则会在集群中默认启动一个 router。
slime 会用 vllm-router 来管理训练过程中的 vLLM 引擎。可以通过--vllm-router-ip与--vllm-router-port来配置 router 的地址。如果不进行配置,则会在集群中默认启动一个 router。
所有的 sglang server 在启动后,会通过/add_worker申请加入 router。在实际进行数据生成的时候,只需要向 router 发送 http 请求,router 会进行 load balancing 操作,将请求转发给 server 们。
所有的 vLLM 引擎在启动后会注册到 router。在实际进行数据生成的时候,只需要向 router 发送 http 请求,router 会进行 load balancing 操作,将请求转发给引擎。
slime 会用 vllm-router 来管理训练过程中的 vLLM 引擎。可以通过--router-ip与--router-port来配置 router 的地址。如果不进行配置,则会在集群中默认启动一个 router。
所有的 vLLM 引擎在启动后会注册到 router。在实际进行数据生成的时候,只需要向 router 发送 http 请求,router 会进行 load balancing 操作,将请求转发给引擎。
当通过--router-ip与--router-port来配置传入一个外部的 router,此时 slime 不再会在内部启动一个 router,而是会把所有的引擎都注册在这个外部 router 上。这时可以利用这个外部的 router 地址来实现更复杂的数据生成流程。注意 router 是支持 openai compatible api 的。
Unable to apply suggestions on deleted lines.
Make sense.
Follow-up to PR #38 review: 1. on-policy-distillation.md (en+zh): `--rm-url` example was `http://<TEACHER_IP>:<TEACHER_PORT>/v1/completions`. Updated to `/inference/v1/generate` — vime's vLLM rollout uses this endpoint by default; the teacher should match. 2. Removed SGLang-comparison advisory blocks. Treating SGLang as if it never existed in vime: - docs/{en,zh}/examples/deepseek-r1.md: - Drop⚠️ "originally written for SGLang backend, check VLLM_ARGS against arguments.py" admonition - Drop "The original SGLang version of this example used large-EP inference (EP64, DP attention, DeepEP); confirm flag availability" - docs/{en,zh}/examples/{glm4.7-30B-A3B,glm4.7-355B-A32B,qwen3-30B-A3B}.md: Drop references to "SGLang's --enable-dp-attention" / "sglang `--enable-dp-attention`". The doc now just describes the vLLM configuration (DP on attention + EP on experts) without invoking SGLang as a reference point. - docs/{en,zh}/blogs/introducing_slime.md: - Drop parenthetical "(originally as an SGLang-native framework; this fork has been adapted to vLLM)" in the article subtitle - Rewrite the "RL-specific workloads" paragraph from past-tense "the original SGLang version upstreamed patches; vLLM port replays..." to present-tense "slime builds on vLLM's /inference/v1/generate ... to support MoE weight updates etc." 3. Final grep result: `grep -E 'sgl|SGL' docs/` returns matches only in docs/{en,zh}/blogs/release_v0.1.0.md (the historical slime v0.1.0 release blog with 6 SGLang mentions each, kept pending separate decision on whether to remove the blog entirely). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Per user direction "当作sglang完全不存在" — remove docs/{en,zh}/blogs/
in its entirety, plus orphan toctree entries:
Removed files (6):
- docs/en/blogs/release_v0.1.0.md (slime upstream v0.1.0 release notes;
sourced from zhihu; tightly coupled to
SGLang-era engineering, 6 sgl mentions)
- docs/zh/blogs/release_v0.1.0.md (same, Chinese)
- docs/en/blogs/introducing_slime.md (vision blog from lmsys.org 2025-07;
SGLang-era origin doc, fully washed
earlier but content overlaps with
README, kept thin value-add)
- docs/zh/blogs/introducing_slime.md (same, Chinese)
- docs/_static/image/blogs/release_v0.1.0/cuda_vmm.png
- docs/_static/image/blogs/release_v0.1.0/overrall.png
Toctree cleanup (2 files):
- docs/en/index.rst:
* Drop "Hardware Platforms" toctree section — its only entry was
`platform_support/amd_tutorial.md`, which #36 deleted from the
tree but did not remove from the toctree (stale reference left
behind on main).
* Drop "Blogs" toctree section (now empty).
- docs/zh/index.rst:
* Drop "博客" toctree section (now empty). (zh side never had the
Hardware Platforms section.)
Verification:
$ git grep -E 'sgl|SGL' -- docs/ # returns nothing — docs/ is now
# 100% free of SGLang references.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…ch.png PR #28 (commit 0d249f9) updated `imgs/arch.png` to the vime architecture diagram. Vime had a second copy at `docs/_static/image/arch.png` referenced by Sphinx docs, which is inherently duplication of the same diagram at two paths. This PR's original diff modified the docs/_static/image/ copy (carried over from PR #18 upstream), but that perpetuates the duplication. Fix: delete the docs/_static/image/arch.png copy and have the two remaining references (docs/{en,zh}/advanced/vllm-config.md) point at the single canonical `imgs/arch.png` via relative path `../../../imgs/`. Single source of truth: 1 arch.png in the tree, used by README, README_zh, and both vllm-config.md docs. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…fig.png topology) The previous regenerate (8b3d1e3) had two wiring bugs vs the original upstream sglang_config.png: 1. MISSING: data buffer → megatron (data feeds training) 2. EXTRA: custom rollout generation → megatron (rollout shouldn't feed training data; that's the data buffer's job) It also had a visual bug — the 'prefill server' / 'decode server' / 'regular server' text labels were centred on the box and overlapping the GPU sprites inside. This commit re-runs matplotlib to produce the topologically correct graph: data buffer → megatron on the left column, custom rollout generation branching to the two server groups, and megatron → prefill server as the weight-sync edge into server group 1. Inner-server labels are now placed above the GPUs. Topology after this commit: data buffer ↔ custom rollout generation │ │ │ ↓ ↓ ↓ megatron ──────→ [server group 1] [server group 2] (vllm router) (vllm router) prefill+decode regular Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…ples (#39) * examples/scripts/slime_plugins: complete sglang→vllm rename + drop SGLang-only examples Split from PR #18. 68 files / +473 / -2180. What this PR does: (A) `examples/` (35→24 files after deletions) and `scripts/` (25 files): rename `--sglang-*` CLI flags to `--vllm-*` equivalents, rename `SGLANG_ARGS=(...)` → `VLLM_ARGS=(...)`, replace `sgl_router_ip` / `sgl_router_port` env var references with the `vllm_router_*` versions, update HF docker image refs, and similar mechanical renames across all training recipes (geo3k_vlm{,_multi_turn} including NPU variants, multi_agent, on_policy_distillation, retool, search-r1, train_infer_mismatch_helper) plus all top-level scripts/run-*.sh and scripts/low_precision/*.sh. (B) `slime_plugins/rollout_buffer/rollout_buffer_example.{py,sh}`: `args.sglang_router_{ip,port}` → `args.vllm_router_{ip,port}`, `pkill -9 sglang` → `pkill -9 vllm`. These are usage examples for the rollout-buffer plugin, not plugin core. (C) Delete SGLang-coupled examples (cannot run on vime, which is vLLM-only after #18 core): - `examples/strands_sglang/` (4 files): entire example targets the `strands-sglang` library (SGLang extension for the `strands` agentic scaffolding). No vLLM-equivalent shim exists upstream; re-introducing this example would require porting strands-sglang itself. - `examples/tau-bench/` (7 files): tau-bench's tool-call parser chain is `trainable_agents.py → openai_tool_adapter.py → sglang_tool_parser.py → sglang.srt.function_call.FunctionCallParser`. With sglang removed from the runtime, the import chain breaks at load time. PR #18 left `sglang_tool_parser.py` untouched and kept the `_sglang`-suffixed method names, so the example never actually worked on vime — the cleanest cut is full removal until someone wants to port to vLLM's tool-calling API as a separate PR. - `examples/README.md`: drop the two list entries that referenced `strands_sglang` and `tau-bench`. (D) `scripts/run-deepseek-r1.sh`: drop the `/sgl-workspace/nvshmem/install/lib/` path from `LD_LIBRARY_PATH`. This was a path that existed only inside the old SGLang docker image (`slimerl/sglang:dev`); vime ships from `vllm/vllm-openai`, where that directory doesn't exist — the `:`-suffixed path was silently ignored at best and risked masking real nvshmem locations at worst. Sgl residue after this PR (3 remaining): - `examples/search-r1/local_dense_retriever/{retrieval_server,download}.py` "Copyright 2023-2024 SGLang Team" headers — legitimate attribution for vendored code, kept. - `examples/retool/README.md` references HuggingFace model id `font-info/qwen3-4b-sft-SGLang-RL` — third-party identifier we cannot change. This PR is independent of the other splits (#38 docs, upcoming tests+CI, PR #18 core). Can land at any time. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * examples/scripts: restore flags dropped by mechanical sglang→vllm rename Reviewing PR #39 against the §6/§7 parameter-mirror table in pr18-post-fix-audit.md, three scripts still had flags that were either dropped silently during the mechanical sed pass, or had comment blocks mangled by the same pass. Mirror table (carry over from the audit): --sglang-cuda-graph-bs LIST → --vllm-cudagraph-capture-sizes LIST --sglang-ep-size N → --vllm-enable-expert-parallel (size auto-derived) --sglang-speculative-{algorithm,num-steps,eagle-topk,num-draft-tokens} → --vllm-speculative-config '{"method":"eagle","num_speculative_tokens":N}' Restored flags: (1) examples/geo3k_vlm/run_geo3k_qwen35.sh: + --vllm-enable-expert-parallel + --vllm-cudagraph-capture-sizes 1 2 4 8 $(seq 16 8 256) + --vllm-speculative-config '{"method":"eagle","num_speculative_tokens":3}' The MTP speculative-decoding block was orphan-commented but had no replacement flag (the 4 spec-decode sglang flags were dropped without the unified config). (2) examples/multi_agent/run-qwen3-30B-A3B-multi-agent.sh: + --vllm-cudagraph-capture-sizes 1 2 4 8 $(seq 16 8 256) The original SGLANG_ARGS had cuda-graph-bs; the new VLLM_ARGS only kept gpu-memory-utilization, dropping cudagraph capture sizes. (3) scripts/low_precision/run-kimi-k2-Thinking-int4.sh: Comment block restoration. The mechanical pass produced corrupted nested comments (`# # --vllm-data-parallel-size 8 / # # # / #`) that lost the original "uncomment to enable DP-attention" intent. Replaced with clear "uncomment to enable" prompts for the two sglang-only-but-now-vllm-flag features (DP-attention via --vllm-data-parallel-size and DeepEP MoE via --vllm-all2all-backend). Verification: re-ran the rename-aware diff scan of every script — all SGLANG_ARGS items from main now have a corresponding VLLM_ARGS item (or are documented sglang-only and intentionally dropped). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * examples/tau-bench: port to vLLM-native tool calling, restore example Reverts the deletion of examples/tau-bench/ from PR #39 commit ad03295 and re-implements the SGLang-coupled tool-call parsing using vLLM's Hermes/Qwen-2.5 tool-call format (as per user direction: "用 vLLM 自带的 tool-calling就行,看vllm源码"). Changes vs the pre-#18 baseline: 1. `examples/tau-bench/sglang_tool_parser.py` → DELETED. It imported `from sglang.srt.function_call.function_call_parser import FunctionCallParser` and `from sglang.srt.managers.io_struct import Function, Tool`. With sglang removed from the vime runtime, this import broke `examples.tau-bench` at load time. 2. `examples/tau-bench/tool_parser.py` → NEW (self-contained, no external sglang/vllm import at runtime). Mirrors vLLM's `Hermes2ProToolParser.extract_tool_calls` (`reference/vllm/vllm/tool_parsers/hermes_tool_parser.py`) for the non-streaming case. Same regex `<tool_call>(.*?)</tool_call>|<tool_call>(.*)` (DOTALL), same JSON-payload shape (`{"name": ..., "arguments": ...}`), same Qwen-2.5 chat template output. Public surface preserves `parse_tools(response, tools, parser)` returning `{"normal_text": str, "calls": [{"name", "parameters"}]}`, so downstream `openai_tool_adapter.py` is untouched in shape. Smoke-tested locally: no tool call, single tool call, malformed JSON (swallowed), multiple tool calls, invalid parser id. 3. `examples/tau-bench/openai_tool_adapter.py`: - Import: `from sglang_tool_parser` → `from tool_parser`. - Rename `_call_to_action_sglang` → `_call_to_action`. - Updated docstrings/comments to drop sglang references. 4. `examples/tau-bench/trainable_agents.py`: - `from slime.rollout.sglang_rollout import GenerateState` → `from slime.rollout.vllm_rollout import GenerateState` (vime moved `GenerateState` to the vllm rollout module). - URL: `http://{sglang_router_ip}:{port}/generate` → `http://{vllm_router_ip}:{port}/v1/completions`. - Payload: `{"text", "sampling_params": {...}}` → `{"prompt", **sampling_params}` (OpenAI completions shape). - Response parsing: `output["text"]` / `output["meta_info"][ "finish_reason"]["type"] == "abort"` → `choice = output[ "choices"][0]; choice["text"]` with abort detection on any `finish_reason` outside `{None, "stop", "length", "tool_calls"}`. - `rollout_args` defaults: `sglang_router_*` → `vllm_router_*`. - `sampling_params` default `max_new_tokens` → `max_tokens` (OpenAI naming). - Rename free function `call_to_action_sglang` → `call_to_action`. - Docstring/comment cleanups dropping sglang references (kept the verl `sglang_multiturn` URL — that's a path on verl's docs site, a third-party identifier). 5. `examples/tau-bench/run_qwen3_4B.sh`: - `pkill -9 sglang` → `pkill -9 -f "vllm serve"`. - `SGLANG_ARGS=(--sglang-mem-fraction-static 0.7 ...)` → `VLLM_ARGS=(--vllm-gpu-memory-utilization 0.7 ...)`. 6. `examples/README.md`: restore the tau-bench list entry that ad03295 removed. Verification: - `grep -E "sglang|sgl_" examples/tau-bench/` returns no real hits (only a verl docs URL containing "sglang_multiturn" as a path, legitimate third-party reference). - `python3 -c "import ast; ast.parse(open('examples/tau-bench/ trainable_agents.py').read())"` → OK. - Local smoke test on `tool_parser.py` covers 5 cases (no call, single call, malformed, multi-call, bad parser id). End-to-end runtime correctness on a test machine is the caller's responsibility — the runtime payload/response shape is a best-effort port; if `vllm-router` surfaces an abort condition differently from what this code assumes, that branch may need adjustment. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * examples/tau-bench: rename tool_parser.py → vllm_tool_parser.py Adds a 'vllm' prefix so the filename signals where the parser was ported from (vLLM's Hermes2ProToolParser) and avoids the ambiguous generic name. Updated callers: - openai_tool_adapter.py: 'from tool_parser' → 'from vllm_tool_parser' - trainable_agents.py docstring: 'tool_parser.parse_tools' → 'vllm_tool_parser.parse_tools' Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * examples/tau-bench: drop entire example After porting tau-bench to vLLM-native tool calling (91ba909 + 85ba197), the user opted to remove the example outright rather than maintain a tau-bench port. The justification: - vime is not actively running tau-bench in CI; the example's true test surface is local-only on the contributor's machine. - openai_tool_adapter.py was ~85% dead code with only a 1-line shim around vllm_tool_parser.parse_tools actually exercised by callers — once dead code is removed, the example is too thin to justify carrying. - Anyone who wants tau-bench can recover it from git history (last present at commit 91ba909) and port further from there. Deleted (7 files): - examples/tau-bench/README.md - examples/tau-bench/generate_with_tau.py - examples/tau-bench/openai_tool_adapter.py - examples/tau-bench/run_qwen3_4B.sh - examples/tau-bench/tau1_mock.py - examples/tau-bench/trainable_agents.py - examples/tau-bench/vllm_tool_parser.py Also dropped the tau-bench list entry from examples/README.md. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Split from PR #18 (gcl/clean-sglang). One of 4 PRs splitting the original PR #18 by content area: docs (#38) / examples (#39) / **tests+CI** / core runtime. 42 files / +~750 / -~700. These are bundled in a single PR because the CI workflows reference test file names by string — splitting them would create a window where either tests are renamed but CI still points at the old names, or vice versa, breaking CI mid-roll. What this PR does: (A) tests/ (38 files): - Mechanical CLI-flag rename: --sglang-* → --vllm-* equivalents in all test scripts (matches the table now used in scripts/ and examples/). - Variable rename: SGLANG_ARGS → VLLM_ARGS where present. - 4 file renames (R086-R091, all >85% similarity): test_qwen2.5_0.5B_opd_sglang.py → test_qwen2.5_0.5B_opd_vllm.py test_qwen2.5_0.5B_sglang_config.py → test_qwen2.5_0.5B_vllm_config.py test_qwen2.5_0.5B_sglang_config_distributed.py → test_qwen2.5_0.5B_vllm_config_distributed.py test_sglang_config_mixed_offload.py → test_vllm_config_mixed_offload.py test_sglang_config_mixed_offload_ft.py → test_vllm_config_mixed_offload_ft.py tests/utils/test_sglang_config.py → tests/utils/test_vllm_config.py - 2 new tests for the IPC weight-transfer path landed in PR #18: tests/test_update_weight_from_tensor.py tests/unit/backends/megatron_utils/update_weight/test_update_weight_from_tensor.py (These are PR #22 / colocate-IPC test coverage; the production code the slim PR #18 ships will rely on the same code from PR #22.) (B) .github/ (4 files): - workflows/conda-ci.yml: container image lmsysorg/sglang → vime (inferactinc/public:vime-vllm-cu129-latest). - workflows/pr-test.yml + pr-test.yml.j2 (template): * Container images (slimerl/slime[-test]:latest → vime image) on every job that ran on the sglang-era base. * e2e-test-sglang-config job → e2e-test-vllm-config job (renamed label `run-ci-sglang-config` → `run-ci-vllm-config`; matrix `test_file` entries updated to point at the renamed test files in (A)). * e2e-test-megatron + e2e-test-image matrices: `_opd_sglang.py` entries → `_opd_vllm.py`. - ISSUE_TEMPLATE/bug_report.yml: drop the "SGLang version (if relevant):" environment field, add "vLLM version:" and "vllm-router version:" lines. (PR #36 already changed "CUDA/ROCm version" → "CUDA version" earlier; that change is preserved.) Sgl residue intentionally kept (4 hits — all anti-regression assertions that prove sglang code paths are gone, not residual references to bring back): - tests/test_update_weight_from_tensor.py:753 — comment "The vLLM IPC implementation must NOT contain sglang-style Gloo gather code". - tests/unit/backends/vllm_utils/test_arguments.py:233-237 — three assertions that --sglang-router-ip, --sglang-router-port, and sglang_router_ip are NOT present in the argument parser. Tests + CI must land together; splitting them risks a window where the CI matrix references test files by names that don't exist yet (or no longer exist). After this lands, the test_file string in CI matches the test files on disk. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: Canlin Guo <canlinguosdu@gmail.com>
There was a problem hiding this comment.
The image is drawn incorrectly; just use the image in /image.
There was a problem hiding this comment.
Error in the drawn image.
* tests + CI: complete sglang→vllm rename across tests/ and .github/ Split from PR #18 (gcl/clean-sglang). One of 4 PRs splitting the original PR #18 by content area: docs (#38) / examples (#39) / **tests+CI** / core runtime. 42 files / +~750 / -~700. These are bundled in a single PR because the CI workflows reference test file names by string — splitting them would create a window where either tests are renamed but CI still points at the old names, or vice versa, breaking CI mid-roll. What this PR does: (A) tests/ (38 files): - Mechanical CLI-flag rename: --sglang-* → --vllm-* equivalents in all test scripts (matches the table now used in scripts/ and examples/). - Variable rename: SGLANG_ARGS → VLLM_ARGS where present. - 4 file renames (R086-R091, all >85% similarity): test_qwen2.5_0.5B_opd_sglang.py → test_qwen2.5_0.5B_opd_vllm.py test_qwen2.5_0.5B_sglang_config.py → test_qwen2.5_0.5B_vllm_config.py test_qwen2.5_0.5B_sglang_config_distributed.py → test_qwen2.5_0.5B_vllm_config_distributed.py test_sglang_config_mixed_offload.py → test_vllm_config_mixed_offload.py test_sglang_config_mixed_offload_ft.py → test_vllm_config_mixed_offload_ft.py tests/utils/test_sglang_config.py → tests/utils/test_vllm_config.py - 2 new tests for the IPC weight-transfer path landed in PR #18: tests/test_update_weight_from_tensor.py tests/unit/backends/megatron_utils/update_weight/test_update_weight_from_tensor.py (These are PR #22 / colocate-IPC test coverage; the production code the slim PR #18 ships will rely on the same code from PR #22.) (B) .github/ (4 files): - workflows/conda-ci.yml: container image lmsysorg/sglang → vime (inferactinc/public:vime-vllm-cu129-latest). - workflows/pr-test.yml + pr-test.yml.j2 (template): * Container images (slimerl/slime[-test]:latest → vime image) on every job that ran on the sglang-era base. * e2e-test-sglang-config job → e2e-test-vllm-config job (renamed label `run-ci-sglang-config` → `run-ci-vllm-config`; matrix `test_file` entries updated to point at the renamed test files in (A)). * e2e-test-megatron + e2e-test-image matrices: `_opd_sglang.py` entries → `_opd_vllm.py`. - ISSUE_TEMPLATE/bug_report.yml: drop the "SGLang version (if relevant):" environment field, add "vLLM version:" and "vllm-router version:" lines. (PR #36 already changed "CUDA/ROCm version" → "CUDA version" earlier; that change is preserved.) Sgl residue intentionally kept (4 hits — all anti-regression assertions that prove sglang code paths are gone, not residual references to bring back): - tests/test_update_weight_from_tensor.py:753 — comment "The vLLM IPC implementation must NOT contain sglang-style Gloo gather code". - tests/unit/backends/vllm_utils/test_arguments.py:233-237 — three assertions that --sglang-router-ip, --sglang-router-port, and sglang_router_ip are NOT present in the argument parser. Tests + CI must land together; splitting them risks a window where the CI matrix references test files by names that don't exist yet (or no longer exist). After this lands, the test_file string in CI matches the test files on disk. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: Canlin Guo <canlinguosdu@gmail.com> * on_policy_distillation: port from SGLang to vLLM /v1/completions Follow-up on the test rename in this PR: test_qwen2.5_0.5B_opd_sglang.py → test_qwen2.5_0.5B_opd_vllm.py. The test only spawns a vLLM teacher and exercises the OPD pipeline; the real broken piece was slime/rollout/on_policy_distillation.py, which PR #18 left in SGLang request/response shape: request fields: "max_new_tokens": 0 (vLLM: "max_tokens") "return_logprob": True (sglang-only) "logprob_start_len": 0 (sglang-only) response parsing: reward["meta_info"]["input_token_logprobs"] (sglang shape) vLLM 0.21 supports the same workflow natively via `prompt_logprobs`: request to POST /v1/completions: { "model": <teacher>, "prompt_token_ids": sample.tokens, "max_tokens": 1, "temperature": 0, "prompt_logprobs": 1, "logprobs": 0, "skip_special_tokens": False, } response: response["choices"][0]["prompt_logprobs"] # list[dict[int, Logprob] | None] where Logprob is {"logprob": float, "rank": int, "decoded_token": str} References checked against vllm source: - reference/vllm/vllm/entrypoints/openai/completion/protocol.py:91 (request: prompt_logprobs: int | None) - reference/vllm/vllm/entrypoints/openai/completion/protocol.py:487 (response: prompt_logprobs: list[dict[int, Logprob] | None] | None) - reference/vllm/vllm/logprobs.py:13 (Logprob dataclass: logprob/rank/decoded_token) Implementation notes: 1. JSON serializes int dict keys as strings, so `_logprob_for_token` tries both `pos_entry.get(token_id)` and `pos_entry.get(str(token_id))`. 2. `pos_entry` is `None` at position 0 (no prior context) — handled explicitly. We also gracefully degrade if a token at position `i` is not in the top-1 logprob dict (falls back to 0.0, same as the prior sglang code would do). 3. The Logprob dataclass `decoded_token` field is unused; we only read `.logprob`. Both dict and `Logprob` shapes are accepted in case the server uses a flatter serialization toggle. 4. `args.opd_teacher_model` is the new model-name arg; falls back to `args.hf_checkpoint` if not set, mirroring how vime's other rollout paths derive the model name. Smoke-tested `_logprob_for_token` locally: - None entry → 0.0 - int key + dict value → logprob - str key (JSON shape) → logprob - missing token → 0.0 - flattened float value → float Also drops 3 lines from tests/unit/backends/vllm_utils/test_arguments.py: the `--sglang-router-ip`/`--sglang-router-port`/`sglang_router_ip` anti- regression assertions. Once the slim PR #18 lands and sglang is gone from the runtime, those assertions are vacuous; treating sglang as non-existent per the project policy. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: Canlin Guo <canlinguosdu@gmail.com> * tests: drop duplicate smoke test updates from PR40 * test(update_weight_from_tensor): drop stale _apply_monkey_patch_torch_reductions patch The inner ``with patch(f"{MODULE_PATH}._apply_monkey_patch_torch_reductions"):`` context in _run_update suppressed a helper call that PR #48 has since deleted from update_weight_from_tensor.py (commit 39bf899 on aoshen/align-ipc-rpc-with-slime). After that PR lands the patched attribute won't exist and this line raises AttributeError. Remove it now so the test survives PR #48 merge. The ``sglang_mod.monkey_patch_torch_reductions = MagicMock()`` stub on the fake sglang module is intentionally kept: on this branch the production code still imports it via ``from ..sglang import monkey_patch_torch_reductions`` (both update_weight_from_tensor._apply_monkey_patch_torch_reductions on PR #40's view of main, and hf_weight_iterator_direct.py at module level). Removing the stub here would break the test on PR #40 alone; it can be dropped in a follow-up once PR #48 finishes removing every import site. Tests: ``tests/unit/backends/megatron_utils/update_weight/test_update_weight_from_tensor.py`` all 6 pass with this change applied to gcl/pr18-tests-ci HEAD. * tests: drop duplicate top-level test_update_weight_from_tensor.py The 786-line tests/test_update_weight_from_tensor.py is a stale rebase leftover from the original PR #18 branch — it predates the IPC test file PR #22 landed at the canonical unit-test path (tests/unit/backends/megatron_utils/update_weight/test_update_weight_from_tensor.py) and predates PR #48's single-RPC weight-version contract. Comparing the two: * Both stub sys.modules / torch.distributed at module import time, so having two files compounds the test-isolation issue Gemini raised (PR #40 comment #1). * Coverage overlaps materially (e.g. test_ipc_init_called_on_first_update_only ≈ test_ipc_init_runs_once — same invariant, different wording). * The nested file is up-to-date with PR #48's RPC contract (update_weights_from_tensor.remote(**fields, weight_version=...)); the top-level file still uses the pre-#48 lifecycle shape and does not exercise the coordinator slot fields. * The nested path matches repo convention: tests/unit/ for mock-only unit tests, tests/ top level for e2e scripts. Closes Gemini comment #1 on PR #40. Gemini comment #2 (the same stub pattern in the surviving nested file) is a pre-existing issue from PR #22 / #48 and out of scope for this rename PR — to be addressed in a follow-up that converts _install_stubs() to an autouse module-scoped fixture with save/restore. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(vllm_config): use real get_model_url default endpoint /inference/v1/generate get_model_url defaults to /inference/v1/generate (PR #18), not /v1/completions. Aligns this test with PR #18's test_vllm_config.py so the two PRs no longer conflict on this file and the assertion matches the actual runtime default. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Drop on_policy_distillation.py from tests+CI PR (now owned by runtime PR #18) The OPD vLLM /v1/completions migration is a runtime change; it was folded into the core-runtime PR (#18). Restore this file to main here so the two PRs no longer overlap on it. #18 merges first, so this lands via #18. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Drop unit-test files now owned by runtime PR #18 test_vllm_config.py + the plugin_contracts tests are coupled to #18's runtime rename (they import vllm_config / vllm_rollout, which #18 creates). They live in #18; remove them here so the two PRs don't overlap. #18 merges first. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test: restore vLLM rollout args dropped during sglang→vllm rename The mechanical sglang→vllm rename dropped several rollout knobs instead of mapping them to their vLLM equivalents, weakening CI coverage (cuda-graph capture caps, speculative decoding, expert parallel). Restore them using the mapping established by the converted production scripts on main (run-glm4.7-30B-A3B.sh / run-glm5-744B-A40B.sh), verified against vLLM AsyncEngineArgs: --sglang-cuda-graph-max-bs N -> --vllm-max-cudagraph-capture-size N --sglang-cuda-graph-bs a b c -> --vllm-cudagraph-capture-sizes a b c --sglang-ep-size N -> --vllm-enable-expert-parallel --sglang-speculative-* (eagle) -> --vllm-speculative-config '{"method":"eagle","num_speculative_tokens":K}' Also: - glm4.7 pd: fix --vllm-max-num-seqs (was 8, taken from cuda-graph-max-bs; --sglang-max-running-requests was 16) and split out cuda-graph capture. - fix sglang→rollout mis-renames in temp-file prefixes (→ vllm_*). - test_vllm_config: rename test_update_weights_default_true → test_update_weights_defaults_to_none (it asserts `is None`). Dropped sglang flags with no vLLM equivalent (enable-dp-lm-head, moe-dense-tp-size, watchdog-timeout, mamba-scheduler-strategy, disaggregation-transfer-backend, enable-metrics) stay dropped; PD KV-transfer is driven by --prefill-num-servers + the --vllm-config prefill/decode topology. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(plugin_contracts): migrate from sglang_rollout to vllm_rollout The three plugin-contract tests still imported slime.rollout.sglang_rollout and called install_stubs(with_sglang_router=True), but _shared.install_stubs already dropped that parameter — so all three failed at collection (TypeError: unexpected keyword 'with_sglang_router'). Complete the migration: - install_stubs(with_sglang_router=True, ...) -> install_stubs(...) - import generate_and_rm / generate_rollout from slime.rollout.vllm_rollout - default rollout/eval path string -> slime.rollout.vllm_rollout.generate_rollout (matches runtime default at slime/utils/arguments.py:233) - FakeGenerateState: sglang_enable_deterministic_inference -> vllm_enable_deterministic_inference, with group_sampling_seeds defaulting to None and gated on the flag (mirrors the already-migrated tests/unit/rollout/test_vllm_rollout.py). All 34 plugin-contract cases pass (were 3 collection errors before). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(update_weight_from_tensor): drop stale slime…megatron_utils.sglang mock The test pre-registered a sys.modules mock for slime.backends.megatron_utils.sglang (monkey_patch_torch_reductions), left over from when update_weight_from_tensor imported it. The module under test no longer imports that module (its real deps are get_gloo_group / HfWeightIteratorBase / update_weight_from_distributed), so the mock is dead. Removing it makes tests/ and .github/ fully sglang-free. Test still passes (7/7). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * [Clean] Remove SGLang runtime code Rebuilt against current main so the PR contains only the SGLang runtime removal -- the docs / tests-ci / examples / scripts / docker portions were split into separate PRs that have since merged. - Delete dead SGLang server/runtime code: sglang_utils/{arguments,sglang_engine}.py, rollout/sglang_rollout.py, the megatron_utils/sglang.py re-export shim, and all docker/**/sglang.patch files. - Rename the rollout config module sglang_utils/sglang_config.py -> vllm_utils/vllm_config.py (SglangConfig -> VllmConfig, _resolve_sglang_config -> _resolve_vllm_config, --sglang-config -> --vllm-config); inline the GPU_MEMORY_TYPE_* constants in rollout.py. - Add megatron_utils/fp8_helpers.py for the UE8M0 fp8 helpers formerly re-exported through the sglang shim; repoint quantizer_fp8 to it. - Swap sglang_router -> vllm_router in http_utils/wandb_utils; drop the dead sglang-router dependency from requirements.txt. - Finish the SGLang->vLLM rename in the runtime so it is internally consistent and matches the tests landing in the tests/CI PR: * router args --router-* -> --vllm-router-* (vllm_router_ip/port/timeout); * get_model_url reads vllm_model_routers (aligning with rollout.py); * --opd-type sglang -> vllm; engine_overrides rename; * sglang_enable_deterministic_inference -> vllm_enable_deterministic_inference, wired to a real --vllm-enable-deterministic-inference flag (exports VLLM_BATCH_INVARIANT=1); * consistent_hash session-id routing uses vllm-router's x-session-id header; * drop dead trace helper build_sglang_meta_trace_attrs; de-SGLang comments/docstrings. - Rename test_sglang_config.py -> test_vllm_config.py and de-SGLang the plugin-contract tests. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Address review: finish de-SGLang + fold OPD/router-policy into runtime - naming: replace residual generic "rollout engine"/"engine"/"comm" wording with concrete vLLM (engine_overrides -> vllm_overrides; arguments help text; http_utils comments; rollout.py "inference workers"). sglang->vllm is correct, sglang->generic is not. - megatron_to_hf: drop the q_a_proj/kv_a_proj_with_mqa pairing + _cached_tensors global. That was sglang-only: sglang's loader torch.cat's both shards within a single load_weights call (needs them co-bucketed), whereas vLLM loads each shard independently via stacked_params_mapping into fused_qkv_a_proj. Also fix the misleading "merge into single fused name" comment. - docker/Dockerfile: remove now-dead sglang/sglang-router --no-deps stubs + the build-time `import sglang` smoke check (slime no longer imports sglang_router). - OPD: migrate on_policy_distillation.py teacher logprobs to vLLM /v1/completions (prompt_logprobs) instead of sglang return_logprob / meta_info.input_token_logprobs. - routing replay: register --vllm-router-policy (dest=router_policy) so the consistent_hash x-session-id session-affinity path is actually wired (was dead). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Review follow-ups: mirror slime vllm_config parsing + restore vLLM process cleanup - vllm_config.from_yaml: drop the needless `models_raw` intermediate and iterate `data["vllm"]` directly, restoring the "Accept both server_groups / legacy engine_groups" comment -- mirrors slime's sglang_config.from_yaml line-for-line. - command_utils.execute_train: re-add a process kill for leftover rollout engines as `pkill -9 -f "vllm serve"` (the old `pkill -9 sglang` was dropped with no vLLM equivalent), so stale engines don't hold GPUs/ports across runs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Drop test changes from runtime PR; tests live in the tests+CI PR (#40) The plugin_contracts tests and the test_sglang_config -> test_vllm_config rename are coupled to the test/CI rename effort and are owned by #40. Restore them to main here so #18 is purely the SGLang runtime removal. #18 merges first; #40 rebases and re-lands the vLLM test versions. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fp8_helpers: copy SGLang verbatim (fix import crash); pin vLLM deep_gemm env - fp8_helpers.py: replace the bespoke rewrite with SGLang's exact implementations of quant_weight_ue8m0 / transform_scale_ue8m0 and their DeepGEMM helpers (per_block_cast_to_fp8, ceil_to_ue8m0, ceil_div, ceil_align, the torch-impl packer). deep_gemm is imported lazily inside the functions (as SGLang does), so module import no longer requires deep_gemm. This fixes the module-level `NameError: _get_tma_aligned_size` that crashed `import megatron_to_hf` on any deep_gemm image, and drops the invented sf-stride fixup block that was not in upstream. Only should_deepgemm_weight_requant_ue8m0 stays vLLM-adapted (is_deep_gemm_e8m0_used) since SGLang's reads SGLang-internal deep_gemm_wrapper. - vllm_engine.launch_server_process: set VLLM_USE_DEEP_GEMM=1 + VLLM_DEEP_GEMM_WARMUP=relax explicitly (setdefault) alongside VLLM_BATCH_INVARIANT, replacing SGLang's removed deep_gemm precompile/warmup envs. All vLLM engine env now lives in the subprocess env builder (single source of truth). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fp8_helpers: revert to vLLM impl + fix the import NameError; add vLLM trace attrs - fp8_helpers.py: keep the vLLM-based implementation (uses vllm.utils.deep_gemm, consistent with the vLLM runtime) rather than the SGLang verbatim copy. Fix the module-level crash: the `try` block referenced `_get_tma_aligned_size` before it was bound (the "pre-imported with fallback" import was never written), which raised NameError whenever deep_gemm imported successfully -- and NameError is not caught by `except ImportError`, so `import megatron_to_hf` crashed on any deep_gemm image. Replace the bogus self-assignment with the real import: `from vllm.utils.deep_gemm import get_tma_aligned_size as _get_tma_aligned_size`. - trace_utils/vllm_rollout: add build_vllm_meta_trace_attrs and attach finish_reason + token usage to the vllm_inference_generate span (mirrors SGLang's build_sglang_meta_trace_attrs; vLLM responses lack the pd_* timing, which lives in vLLM's own OTLP traces). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(opd): score teacher via /inference/v1/generate with prompt_logprobs Move the vllm OPD teacher path off the OpenAI /v1/completions endpoint onto vime's native /inference/v1/generate (the same endpoint the rollout engines use), and fix three latent issues: 1. model field: /inference/v1/generate takes `model` as OPTIONAL. Stop defaulting to args.hf_checkpoint (the *student* name, which mis-names a teacher!=student server). Add --opd-teacher-model; send `model` only when set, otherwise omit it (single-model teacher servers use their loaded model). 2. multimodal: the old code sent image_data to a token-only endpoint, which is invalid. Raise NotImplementedError until the /v1/chat/completions/render -> /inference/v1/generate flow is wired (mirrors slime.rollout.vllm_rollout.generate). 3. logprob robustness: read top-level GenerateResponse.prompt_logprobs, assert it is present and length-aligned with token_ids, assert the per-sample tensor covers response_length, and raise (not silently return 0.0) on a missing token logprob. vLLM always includes the actual prompt token in prompt_logprobs, so a miss is a real error. Alignment is unchanged (plp[i] <-> tokens[i], skip pos 0, take [-response_length:]). Follow-up (separate, in the tests PR): the OPD e2e test must launch a teacher that exposes /inference/v1/generate and point --rm-url at it. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(clean-sglang): purge SGLang from tools, train scripts, and build infra tools/: drop dead `args.sglang_enable_ep_moe` shim (read nowhere); reword profile/replay helpers to vLLM and map analyzer hints to vLLM flags (--enforce-eager, --gpu-memory-utilization). train{,_async}.py: comments SGLang -> vLLM. build infra: remove build_conda.sh (SGLang-only conda path); drop the GB300 sgl-kernel install from the Dockerfile; delete docker/npu_patch/ wholesale. docker base image: bump to vLLM v0.22.0. justfile ARM recipes now pin the real multi-arch vLLM base images instead of the dead SGLANG_IMAGE_TAG/ ENABLE_SGLANG_PATCH build-args -- cu129-arm64 -> v0.22.0-cu129-ubuntu2404 (CUDA 12.9), cu13-arm64 -> v0.22.0-ubuntu2404 (the default-CUDA tag is already CUDA 13.0) + ENABLE_CUDA_13=1. vLLM tags are multi-arch manifests, so docker selects the arm64 image automatically on an ARM host. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(clean-sglang): drop conda-build workflow (ran deleted build_conda.sh on SGLang image) The single build-conda job ran `bash build_conda.sh` (removed in the previous commit) inside an lmsysorg/sglang container. With the SGLang-only conda path gone, the whole workflow is dead. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(clean-sglang): fix stale SGLang refs in docs/skills; tidy comments docs/conf.py: point the "edit on GitHub" links at vllm-project/vime instead of the inherited sgl-project.github.io repo. .claude/skills/*: update the dead `slime/rollout/sglang_rollout.py` references to `vllm_rollout.py` (the real default is slime.rollout.vllm_rollout.generate_rollout). justfile: drop the redundant BASE_IMAGE override on release-cu129-arm64 (it equalled the Dockerfile default; the multi-arch manifest already resolves arm64). train{,_async}.py: drop stray "the" in the W&B comment. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(tests): use method=mtp (not eagle) in vllm speculative config The migrated speculative configs pass no draft `model`, so method=eagle raises "num_speculative_tokens was provided but without speculative model" in vLLM's SpeculativeConfig. These models carry embedded MTP layers, so method=mtp is correct and unblocks the mimo MTP-only-grad test (#19). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(cleanup): target renamed vLLM subprocesses in pkill so VRAM is freed vLLM's set_process_title() renames the VRAM-holding subprocesses (VLLM::EngineCore, VLLM::Worker_TP*, vllm::router), so their cmdline no longer contains "vllm serve". The previous `pkill -9 -f "vllm serve"` matched only the launcher and left engine/worker children holding GPU memory, leaking it into the next run — masked only by the indiscriminate `pkill -9 python`, which is unsafe on colocate/shared nodes. Match both the launcher and the renamed children with `pkill -9 -f '[v]llm serve|VLL[M]::'`; the [v]/[M] bracket trick keeps the pattern from matching pkill's own cmdline. This makes the broad python kill unnecessary, so its already-commented-out lines are removed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(speculative): use method=mtp (not eagle) for embedded-MTP models vLLM's SpeculativeConfig requires an explicit draft `model` for method=eagle; with only num_speculative_tokens set it raises "num_speculative_tokens was provided but without speculative model". The migrated configs in scripts/examples/docs pass no model, so they must use method=mtp, which reuses the target checkpoint's embedded MTP layer (DeepSeek-R1, GLM-4.x-MoE, MiMo, Qwen3-Next/3.5). The two docs examples that pass an explicit "model" are genuine eagle usage and are left unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(vllm): launch each rollout engine with its ServerGroup's per-group TP launch_server_process / _init_normal derived tensor-parallel size and CUDA_VISIBLE_DEVICES from the global --rollout-num-gpus-per-engine, ignoring the per-engine num_gpus_per_engine already carried on the VLLMEngine actor. A ServerGroup configured with num_gpus_per_engine greater than the global flag (e.g. tp=2) therefore launched as tp=1, while the NCCL weight-sync rendezvous sized world_size from engine_gpu_counts (the per-group value). The two disagreed: the trainer waited for a rank the under-sized engine never started, so init_weight_transfer_engine hung for 300s ("3/4 clients joined") and the job failed. Honor the per-engine num_gpus_per_engine at launch, falling back to the global flag when unset (matches the SGLang path and PR #66's _compute_server_args). Verified on H200: tests/test_qwen2.5_0.5B_vllm_config_distributed now launches engine0 tp=2 / engine1 tp=1, update_weights completes in 1.1s (was a 301s timeout), and rollout+eval proceed. AI assistance (Claude Code) was used for this change. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(ckpt): add --dist-ckpt-optim-fully-reshardable for PAO+offload save/load test_qwen3_4B_ckpt.py uses precision-aware optimizer + cpu-offload (HybridDeviceOptimizer). Under the default dp_reshardable (bucket-centric) optimizer sharding, save/load produce unequal-length param_state lists, so dist-ckpt load fails with "Cannot merge two lists with different lengths (81 and 79)". fully_reshardable is model-centric and immune to bucket-layout changes. Verified on the r3 image (Megatron-LM 0.16.0rc0 @ 1dcf0da): save+load both succeed, and source review confirms master_param / step / HybridDeviceOptimizer sync are handled on this path. This is the flag described in PR #50 that was never actually merged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * refactor(router-args): hybrid naming — vllm_ for ip/port, bare router_ for timeout vllm-router's RouterArgs.from_cli_args only supports prefix "" or "router_" (never "vllm_router_"), and excludes host/port from its CLI via exclude_host_port=True. So: - --vllm-router-ip / --vllm-router-port keep the vllm_ prefix: RouterArgs does not own these CLI flags, vime does (populated via _start_router's manual router_args.host/port assignment), so the vllm_ prefix is free and marks them as vime-owned endpoint config. - --router-request-timeout-secs goes bare (dest router_request_timeout_secs): it is a genuine RouterArgs field, so it shares the --router-* namespace with policy / cache_threshold / retries / … and flows through from_cli_args like the other knobs. - --vllm-router-policy keeps dest=router_policy (unchanged). Also fixes conftest fixture to seed vllm_router_ip/port (was bare router_ip/port, which never matched the vllm_engine reader) and updates README/README_zh prose. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: Canlin Guo <canlinguosdu@gmail.com>
…ples (#39) * examples/scripts/slime_plugins: complete sglang→vllm rename + drop SGLang-only examples Split from PR #18. 68 files / +473 / -2180. What this PR does: (A) `examples/` (35→24 files after deletions) and `scripts/` (25 files): rename `--sglang-*` CLI flags to `--vllm-*` equivalents, rename `SGLANG_ARGS=(...)` → `VLLM_ARGS=(...)`, replace `sgl_router_ip` / `sgl_router_port` env var references with the `vllm_router_*` versions, update HF docker image refs, and similar mechanical renames across all training recipes (geo3k_vlm{,_multi_turn} including NPU variants, multi_agent, on_policy_distillation, retool, search-r1, train_infer_mismatch_helper) plus all top-level scripts/run-*.sh and scripts/low_precision/*.sh. (B) `slime_plugins/rollout_buffer/rollout_buffer_example.{py,sh}`: `args.sglang_router_{ip,port}` → `args.vllm_router_{ip,port}`, `pkill -9 sglang` → `pkill -9 vllm`. These are usage examples for the rollout-buffer plugin, not plugin core. (C) Delete SGLang-coupled examples (cannot run on vime, which is vLLM-only after #18 core): - `examples/strands_sglang/` (4 files): entire example targets the `strands-sglang` library (SGLang extension for the `strands` agentic scaffolding). No vLLM-equivalent shim exists upstream; re-introducing this example would require porting strands-sglang itself. - `examples/tau-bench/` (7 files): tau-bench's tool-call parser chain is `trainable_agents.py → openai_tool_adapter.py → sglang_tool_parser.py → sglang.srt.function_call.FunctionCallParser`. With sglang removed from the runtime, the import chain breaks at load time. PR #18 left `sglang_tool_parser.py` untouched and kept the `_sglang`-suffixed method names, so the example never actually worked on vime — the cleanest cut is full removal until someone wants to port to vLLM's tool-calling API as a separate PR. - `examples/README.md`: drop the two list entries that referenced `strands_sglang` and `tau-bench`. (D) `scripts/run-deepseek-r1.sh`: drop the `/sgl-workspace/nvshmem/install/lib/` path from `LD_LIBRARY_PATH`. This was a path that existed only inside the old SGLang docker image (`slimerl/sglang:dev`); vime ships from `vllm/vllm-openai`, where that directory doesn't exist — the `:`-suffixed path was silently ignored at best and risked masking real nvshmem locations at worst. Sgl residue after this PR (3 remaining): - `examples/search-r1/local_dense_retriever/{retrieval_server,download}.py` "Copyright 2023-2024 SGLang Team" headers — legitimate attribution for vendored code, kept. - `examples/retool/README.md` references HuggingFace model id `font-info/qwen3-4b-sft-SGLang-RL` — third-party identifier we cannot change. This PR is independent of the other splits (#38 docs, upcoming tests+CI, PR #18 core). Can land at any time. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * examples/scripts: restore flags dropped by mechanical sglang→vllm rename Reviewing PR #39 against the §6/§7 parameter-mirror table in pr18-post-fix-audit.md, three scripts still had flags that were either dropped silently during the mechanical sed pass, or had comment blocks mangled by the same pass. Mirror table (carry over from the audit): --sglang-cuda-graph-bs LIST → --vllm-cudagraph-capture-sizes LIST --sglang-ep-size N → --vllm-enable-expert-parallel (size auto-derived) --sglang-speculative-{algorithm,num-steps,eagle-topk,num-draft-tokens} → --vllm-speculative-config '{"method":"eagle","num_speculative_tokens":N}' Restored flags: (1) examples/geo3k_vlm/run_geo3k_qwen35.sh: + --vllm-enable-expert-parallel + --vllm-cudagraph-capture-sizes 1 2 4 8 $(seq 16 8 256) + --vllm-speculative-config '{"method":"eagle","num_speculative_tokens":3}' The MTP speculative-decoding block was orphan-commented but had no replacement flag (the 4 spec-decode sglang flags were dropped without the unified config). (2) examples/multi_agent/run-qwen3-30B-A3B-multi-agent.sh: + --vllm-cudagraph-capture-sizes 1 2 4 8 $(seq 16 8 256) The original SGLANG_ARGS had cuda-graph-bs; the new VLLM_ARGS only kept gpu-memory-utilization, dropping cudagraph capture sizes. (3) scripts/low_precision/run-kimi-k2-Thinking-int4.sh: Comment block restoration. The mechanical pass produced corrupted nested comments (`# # --vllm-data-parallel-size 8 / # # # / #`) that lost the original "uncomment to enable DP-attention" intent. Replaced with clear "uncomment to enable" prompts for the two sglang-only-but-now-vllm-flag features (DP-attention via --vllm-data-parallel-size and DeepEP MoE via --vllm-all2all-backend). Verification: re-ran the rename-aware diff scan of every script — all SGLANG_ARGS items from main now have a corresponding VLLM_ARGS item (or are documented sglang-only and intentionally dropped). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * examples/tau-bench: port to vLLM-native tool calling, restore example Reverts the deletion of examples/tau-bench/ from PR #39 commit ad03295 and re-implements the SGLang-coupled tool-call parsing using vLLM's Hermes/Qwen-2.5 tool-call format (as per user direction: "用 vLLM 自带的 tool-calling就行,看vllm源码"). Changes vs the pre-#18 baseline: 1. `examples/tau-bench/sglang_tool_parser.py` → DELETED. It imported `from sglang.srt.function_call.function_call_parser import FunctionCallParser` and `from sglang.srt.managers.io_struct import Function, Tool`. With sglang removed from the vime runtime, this import broke `examples.tau-bench` at load time. 2. `examples/tau-bench/tool_parser.py` → NEW (self-contained, no external sglang/vllm import at runtime). Mirrors vLLM's `Hermes2ProToolParser.extract_tool_calls` (`reference/vllm/vllm/tool_parsers/hermes_tool_parser.py`) for the non-streaming case. Same regex `<tool_call>(.*?)</tool_call>|<tool_call>(.*)` (DOTALL), same JSON-payload shape (`{"name": ..., "arguments": ...}`), same Qwen-2.5 chat template output. Public surface preserves `parse_tools(response, tools, parser)` returning `{"normal_text": str, "calls": [{"name", "parameters"}]}`, so downstream `openai_tool_adapter.py` is untouched in shape. Smoke-tested locally: no tool call, single tool call, malformed JSON (swallowed), multiple tool calls, invalid parser id. 3. `examples/tau-bench/openai_tool_adapter.py`: - Import: `from sglang_tool_parser` → `from tool_parser`. - Rename `_call_to_action_sglang` → `_call_to_action`. - Updated docstrings/comments to drop sglang references. 4. `examples/tau-bench/trainable_agents.py`: - `from slime.rollout.sglang_rollout import GenerateState` → `from slime.rollout.vllm_rollout import GenerateState` (vime moved `GenerateState` to the vllm rollout module). - URL: `http://{sglang_router_ip}:{port}/generate` → `http://{vllm_router_ip}:{port}/v1/completions`. - Payload: `{"text", "sampling_params": {...}}` → `{"prompt", **sampling_params}` (OpenAI completions shape). - Response parsing: `output["text"]` / `output["meta_info"][ "finish_reason"]["type"] == "abort"` → `choice = output[ "choices"][0]; choice["text"]` with abort detection on any `finish_reason` outside `{None, "stop", "length", "tool_calls"}`. - `rollout_args` defaults: `sglang_router_*` → `vllm_router_*`. - `sampling_params` default `max_new_tokens` → `max_tokens` (OpenAI naming). - Rename free function `call_to_action_sglang` → `call_to_action`. - Docstring/comment cleanups dropping sglang references (kept the verl `sglang_multiturn` URL — that's a path on verl's docs site, a third-party identifier). 5. `examples/tau-bench/run_qwen3_4B.sh`: - `pkill -9 sglang` → `pkill -9 -f "vllm serve"`. - `SGLANG_ARGS=(--sglang-mem-fraction-static 0.7 ...)` → `VLLM_ARGS=(--vllm-gpu-memory-utilization 0.7 ...)`. 6. `examples/README.md`: restore the tau-bench list entry that ad03295 removed. Verification: - `grep -E "sglang|sgl_" examples/tau-bench/` returns no real hits (only a verl docs URL containing "sglang_multiturn" as a path, legitimate third-party reference). - `python3 -c "import ast; ast.parse(open('examples/tau-bench/ trainable_agents.py').read())"` → OK. - Local smoke test on `tool_parser.py` covers 5 cases (no call, single call, malformed, multi-call, bad parser id). End-to-end runtime correctness on a test machine is the caller's responsibility — the runtime payload/response shape is a best-effort port; if `vllm-router` surfaces an abort condition differently from what this code assumes, that branch may need adjustment. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * examples/tau-bench: rename tool_parser.py → vllm_tool_parser.py Adds a 'vllm' prefix so the filename signals where the parser was ported from (vLLM's Hermes2ProToolParser) and avoids the ambiguous generic name. Updated callers: - openai_tool_adapter.py: 'from tool_parser' → 'from vllm_tool_parser' - trainable_agents.py docstring: 'tool_parser.parse_tools' → 'vllm_tool_parser.parse_tools' Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * examples/tau-bench: drop entire example After porting tau-bench to vLLM-native tool calling (91ba909 + 85ba197), the user opted to remove the example outright rather than maintain a tau-bench port. The justification: - vime is not actively running tau-bench in CI; the example's true test surface is local-only on the contributor's machine. - openai_tool_adapter.py was ~85% dead code with only a 1-line shim around vllm_tool_parser.parse_tools actually exercised by callers — once dead code is removed, the example is too thin to justify carrying. - Anyone who wants tau-bench can recover it from git history (last present at commit 91ba909) and port further from there. Deleted (7 files): - examples/tau-bench/README.md - examples/tau-bench/generate_with_tau.py - examples/tau-bench/openai_tool_adapter.py - examples/tau-bench/run_qwen3_4B.sh - examples/tau-bench/tau1_mock.py - examples/tau-bench/trainable_agents.py - examples/tau-bench/vllm_tool_parser.py Also dropped the tau-bench list entry from examples/README.md. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs: complete sglang→vllm rename across docs/ + README Split from PR #18 (gcl/clean-sglang). One of 4 PRs splitting the original PR #18 by content area (docs / examples / tests+CI / core runtime). See pr18-split-plan.md. 53 files. Pure documentation: no runtime code change. What this PR does: - docs/en/ + docs/zh/ (49 files): rename sglang-config.md → vllm-config.md (1 each side); update all --sglang-* CLI flag examples to their --vllm-* equivalents; rename references to SGLang as the rollout backend → vLLM; fix three broken upstream URLs (each replacement verified to load): * blog.vllm.ai/2025/10/26/distributed-batch-invariance.html → vllm.ai/blog/2025-11-10-bitwise-consistent-train-inference * docs.vllm.ai/.../features/spec_decode.html (301 → 404) → docs.vllm.ai/.../features/speculative_decoding/ * github.com/vllm-project/vllm-router (404; repo does not exist) → github.com/vllm-project/router (the actual router repo) - docs/_static/image/ (3 files): rename sglang_config.png → vllm_config.png; refresh arch.png to the vllm-router architecture. - docs/en/advanced/reproducibility.md and docs/zh/advanced/reproducibility.md: the en and zh files had their content swapped (en had Chinese body, zh had English body). Restored the correct mapping. - docs/{en,zh}/blogs/introducing_slime.md: remove the "Special thanks to the AMD GenAI - Foundation Model Team for Day-1 AMD hardware support" line. With AMD support removed in #36, the acknowledgement implied ongoing support that no longer exists. - README.md / README_zh.md: expanded vLLM router argument prefix explanation (--router-* for vllm-router native flags vs --vllm-router-* for vime-side orchestration knobs); remove stale link to docs/zh/advanced/ slime_vllm_backend_design_v1.md which was deleted in commit fe700db. This PR is independent of the other splits and can land at any time without coordination. See pr18-split-plan.md §"合并顺序无强约束". Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs: scrub remaining sglang references and fix OPD teacher URL Follow-up to PR #38 review: 1. on-policy-distillation.md (en+zh): `--rm-url` example was `http://<TEACHER_IP>:<TEACHER_PORT>/v1/completions`. Updated to `/inference/v1/generate` — vime's vLLM rollout uses this endpoint by default; the teacher should match. 2. Removed SGLang-comparison advisory blocks. Treating SGLang as if it never existed in vime: - docs/{en,zh}/examples/deepseek-r1.md: - Drop⚠️ "originally written for SGLang backend, check VLLM_ARGS against arguments.py" admonition - Drop "The original SGLang version of this example used large-EP inference (EP64, DP attention, DeepEP); confirm flag availability" - docs/{en,zh}/examples/{glm4.7-30B-A3B,glm4.7-355B-A32B,qwen3-30B-A3B}.md: Drop references to "SGLang's --enable-dp-attention" / "sglang `--enable-dp-attention`". The doc now just describes the vLLM configuration (DP on attention + EP on experts) without invoking SGLang as a reference point. - docs/{en,zh}/blogs/introducing_slime.md: - Drop parenthetical "(originally as an SGLang-native framework; this fork has been adapted to vLLM)" in the article subtitle - Rewrite the "RL-specific workloads" paragraph from past-tense "the original SGLang version upstreamed patches; vLLM port replays..." to present-tense "slime builds on vLLM's /inference/v1/generate ... to support MoE weight updates etc." 3. Final grep result: `grep -E 'sgl|SGL' docs/` returns matches only in docs/{en,zh}/blogs/release_v0.1.0.md (the historical slime v0.1.0 release blog with 6 SGLang mentions each, kept pending separate decision on whether to remove the blog entirely). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs: remove historical slime blog posts + clean up index.rst toctree Per user direction "当作sglang完全不存在" — remove docs/{en,zh}/blogs/ in its entirety, plus orphan toctree entries: Removed files (6): - docs/en/blogs/release_v0.1.0.md (slime upstream v0.1.0 release notes; sourced from zhihu; tightly coupled to SGLang-era engineering, 6 sgl mentions) - docs/zh/blogs/release_v0.1.0.md (same, Chinese) - docs/en/blogs/introducing_slime.md (vision blog from lmsys.org 2025-07; SGLang-era origin doc, fully washed earlier but content overlaps with README, kept thin value-add) - docs/zh/blogs/introducing_slime.md (same, Chinese) - docs/_static/image/blogs/release_v0.1.0/cuda_vmm.png - docs/_static/image/blogs/release_v0.1.0/overrall.png Toctree cleanup (2 files): - docs/en/index.rst: * Drop "Hardware Platforms" toctree section — its only entry was `platform_support/amd_tutorial.md`, which #36 deleted from the tree but did not remove from the toctree (stale reference left behind on main). * Drop "Blogs" toctree section (now empty). - docs/zh/index.rst: * Drop "博客" toctree section (now empty). (zh side never had the Hardware Platforms section.) Verification: $ git grep -E 'sgl|SGL' -- docs/ # returns nothing — docs/ is now # 100% free of SGLang references. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs: dedup arch.png — drop docs/_static/image copy, point to imgs/arch.png PR #28 (commit 0d249f9) updated `imgs/arch.png` to the vime architecture diagram. Vime had a second copy at `docs/_static/image/arch.png` referenced by Sphinx docs, which is inherently duplication of the same diagram at two paths. This PR's original diff modified the docs/_static/image/ copy (carried over from PR #18 upstream), but that perpetuates the duplication. Fix: delete the docs/_static/image/arch.png copy and have the two remaining references (docs/{en,zh}/advanced/vllm-config.md) point at the single canonical `imgs/arch.png` via relative path `../../../imgs/`. Single source of truth: 1 arch.png in the tree, used by README, README_zh, and both vllm-config.md docs. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs: fix vllm_config.png wiring (redraw to match original sglang_config.png topology) The previous regenerate (8b3d1e3) had two wiring bugs vs the original upstream sglang_config.png: 1. MISSING: data buffer → megatron (data feeds training) 2. EXTRA: custom rollout generation → megatron (rollout shouldn't feed training data; that's the data buffer's job) It also had a visual bug — the 'prefill server' / 'decode server' / 'regular server' text labels were centred on the box and overlapping the GPU sprites inside. This commit re-runs matplotlib to produce the topologically correct graph: data buffer → megatron on the left column, custom rollout generation branching to the two server groups, and megatron → prefill server as the weight-sync edge into server group 1. Inner-server labels are now placed above the GPUs. Topology after this commit: data buffer ↔ custom rollout generation │ │ │ ↓ ↓ ↓ megatron ──────→ [server group 1] [server group 2] (vllm router) (vllm router) prefill+decode regular Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs: mention TorchSpec for draft training * update architecture image. --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: kaiyuan <kyxiezju@163.com>
* tests + CI: complete sglang→vllm rename across tests/ and .github/ Split from PR #18 (gcl/clean-sglang). One of 4 PRs splitting the original PR #18 by content area: docs (#38) / examples (#39) / **tests+CI** / core runtime. 42 files / +~750 / -~700. These are bundled in a single PR because the CI workflows reference test file names by string — splitting them would create a window where either tests are renamed but CI still points at the old names, or vice versa, breaking CI mid-roll. What this PR does: (A) tests/ (38 files): - Mechanical CLI-flag rename: --sglang-* → --vllm-* equivalents in all test scripts (matches the table now used in scripts/ and examples/). - Variable rename: SGLANG_ARGS → VLLM_ARGS where present. - 4 file renames (R086-R091, all >85% similarity): test_qwen2.5_0.5B_opd_sglang.py → test_qwen2.5_0.5B_opd_vllm.py test_qwen2.5_0.5B_sglang_config.py → test_qwen2.5_0.5B_vllm_config.py test_qwen2.5_0.5B_sglang_config_distributed.py → test_qwen2.5_0.5B_vllm_config_distributed.py test_sglang_config_mixed_offload.py → test_vllm_config_mixed_offload.py test_sglang_config_mixed_offload_ft.py → test_vllm_config_mixed_offload_ft.py tests/utils/test_sglang_config.py → tests/utils/test_vllm_config.py - 2 new tests for the IPC weight-transfer path landed in PR #18: tests/test_update_weight_from_tensor.py tests/unit/backends/megatron_utils/update_weight/test_update_weight_from_tensor.py (These are PR #22 / colocate-IPC test coverage; the production code the slim PR #18 ships will rely on the same code from PR #22.) (B) .github/ (4 files): - workflows/conda-ci.yml: container image lmsysorg/sglang → vime (inferactinc/public:vime-vllm-cu129-latest). - workflows/pr-test.yml + pr-test.yml.j2 (template): * Container images (slimerl/slime[-test]:latest → vime image) on every job that ran on the sglang-era base. * e2e-test-sglang-config job → e2e-test-vllm-config job (renamed label `run-ci-sglang-config` → `run-ci-vllm-config`; matrix `test_file` entries updated to point at the renamed test files in (A)). * e2e-test-megatron + e2e-test-image matrices: `_opd_sglang.py` entries → `_opd_vllm.py`. - ISSUE_TEMPLATE/bug_report.yml: drop the "SGLang version (if relevant):" environment field, add "vLLM version:" and "vllm-router version:" lines. (PR #36 already changed "CUDA/ROCm version" → "CUDA version" earlier; that change is preserved.) Sgl residue intentionally kept (4 hits — all anti-regression assertions that prove sglang code paths are gone, not residual references to bring back): - tests/test_update_weight_from_tensor.py:753 — comment "The vLLM IPC implementation must NOT contain sglang-style Gloo gather code". - tests/unit/backends/vllm_utils/test_arguments.py:233-237 — three assertions that --sglang-router-ip, --sglang-router-port, and sglang_router_ip are NOT present in the argument parser. Tests + CI must land together; splitting them risks a window where the CI matrix references test files by names that don't exist yet (or no longer exist). After this lands, the test_file string in CI matches the test files on disk. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: Canlin Guo <canlinguosdu@gmail.com> * on_policy_distillation: port from SGLang to vLLM /v1/completions Follow-up on the test rename in this PR: test_qwen2.5_0.5B_opd_sglang.py → test_qwen2.5_0.5B_opd_vllm.py. The test only spawns a vLLM teacher and exercises the OPD pipeline; the real broken piece was slime/rollout/on_policy_distillation.py, which PR #18 left in SGLang request/response shape: request fields: "max_new_tokens": 0 (vLLM: "max_tokens") "return_logprob": True (sglang-only) "logprob_start_len": 0 (sglang-only) response parsing: reward["meta_info"]["input_token_logprobs"] (sglang shape) vLLM 0.21 supports the same workflow natively via `prompt_logprobs`: request to POST /v1/completions: { "model": <teacher>, "prompt_token_ids": sample.tokens, "max_tokens": 1, "temperature": 0, "prompt_logprobs": 1, "logprobs": 0, "skip_special_tokens": False, } response: response["choices"][0]["prompt_logprobs"] # list[dict[int, Logprob] | None] where Logprob is {"logprob": float, "rank": int, "decoded_token": str} References checked against vllm source: - reference/vllm/vllm/entrypoints/openai/completion/protocol.py:91 (request: prompt_logprobs: int | None) - reference/vllm/vllm/entrypoints/openai/completion/protocol.py:487 (response: prompt_logprobs: list[dict[int, Logprob] | None] | None) - reference/vllm/vllm/logprobs.py:13 (Logprob dataclass: logprob/rank/decoded_token) Implementation notes: 1. JSON serializes int dict keys as strings, so `_logprob_for_token` tries both `pos_entry.get(token_id)` and `pos_entry.get(str(token_id))`. 2. `pos_entry` is `None` at position 0 (no prior context) — handled explicitly. We also gracefully degrade if a token at position `i` is not in the top-1 logprob dict (falls back to 0.0, same as the prior sglang code would do). 3. The Logprob dataclass `decoded_token` field is unused; we only read `.logprob`. Both dict and `Logprob` shapes are accepted in case the server uses a flatter serialization toggle. 4. `args.opd_teacher_model` is the new model-name arg; falls back to `args.hf_checkpoint` if not set, mirroring how vime's other rollout paths derive the model name. Smoke-tested `_logprob_for_token` locally: - None entry → 0.0 - int key + dict value → logprob - str key (JSON shape) → logprob - missing token → 0.0 - flattened float value → float Also drops 3 lines from tests/unit/backends/vllm_utils/test_arguments.py: the `--sglang-router-ip`/`--sglang-router-port`/`sglang_router_ip` anti- regression assertions. Once the slim PR #18 lands and sglang is gone from the runtime, those assertions are vacuous; treating sglang as non-existent per the project policy. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: Canlin Guo <canlinguosdu@gmail.com> * tests: drop duplicate smoke test updates from PR40 * test(update_weight_from_tensor): drop stale _apply_monkey_patch_torch_reductions patch The inner ``with patch(f"{MODULE_PATH}._apply_monkey_patch_torch_reductions"):`` context in _run_update suppressed a helper call that PR #48 has since deleted from update_weight_from_tensor.py (commit 39bf899 on aoshen/align-ipc-rpc-with-slime). After that PR lands the patched attribute won't exist and this line raises AttributeError. Remove it now so the test survives PR #48 merge. The ``sglang_mod.monkey_patch_torch_reductions = MagicMock()`` stub on the fake sglang module is intentionally kept: on this branch the production code still imports it via ``from ..sglang import monkey_patch_torch_reductions`` (both update_weight_from_tensor._apply_monkey_patch_torch_reductions on PR #40's view of main, and hf_weight_iterator_direct.py at module level). Removing the stub here would break the test on PR #40 alone; it can be dropped in a follow-up once PR #48 finishes removing every import site. Tests: ``tests/unit/backends/megatron_utils/update_weight/test_update_weight_from_tensor.py`` all 6 pass with this change applied to gcl/pr18-tests-ci HEAD. * tests: drop duplicate top-level test_update_weight_from_tensor.py The 786-line tests/test_update_weight_from_tensor.py is a stale rebase leftover from the original PR #18 branch — it predates the IPC test file PR #22 landed at the canonical unit-test path (tests/unit/backends/megatron_utils/update_weight/test_update_weight_from_tensor.py) and predates PR #48's single-RPC weight-version contract. Comparing the two: * Both stub sys.modules / torch.distributed at module import time, so having two files compounds the test-isolation issue Gemini raised (PR #40 comment #1). * Coverage overlaps materially (e.g. test_ipc_init_called_on_first_update_only ≈ test_ipc_init_runs_once — same invariant, different wording). * The nested file is up-to-date with PR #48's RPC contract (update_weights_from_tensor.remote(**fields, weight_version=...)); the top-level file still uses the pre-#48 lifecycle shape and does not exercise the coordinator slot fields. * The nested path matches repo convention: tests/unit/ for mock-only unit tests, tests/ top level for e2e scripts. Closes Gemini comment #1 on PR #40. Gemini comment #2 (the same stub pattern in the surviving nested file) is a pre-existing issue from PR #22 / #48 and out of scope for this rename PR — to be addressed in a follow-up that converts _install_stubs() to an autouse module-scoped fixture with save/restore. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(vllm_config): use real get_model_url default endpoint /inference/v1/generate get_model_url defaults to /inference/v1/generate (PR #18), not /v1/completions. Aligns this test with PR #18's test_vllm_config.py so the two PRs no longer conflict on this file and the assertion matches the actual runtime default. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Drop on_policy_distillation.py from tests+CI PR (now owned by runtime PR #18) The OPD vLLM /v1/completions migration is a runtime change; it was folded into the core-runtime PR (#18). Restore this file to main here so the two PRs no longer overlap on it. #18 merges first, so this lands via #18. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Drop unit-test files now owned by runtime PR #18 test_vllm_config.py + the plugin_contracts tests are coupled to #18's runtime rename (they import vllm_config / vllm_rollout, which #18 creates). They live in #18; remove them here so the two PRs don't overlap. #18 merges first. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test: restore vLLM rollout args dropped during sglang→vllm rename The mechanical sglang→vllm rename dropped several rollout knobs instead of mapping them to their vLLM equivalents, weakening CI coverage (cuda-graph capture caps, speculative decoding, expert parallel). Restore them using the mapping established by the converted production scripts on main (run-glm4.7-30B-A3B.sh / run-glm5-744B-A40B.sh), verified against vLLM AsyncEngineArgs: --sglang-cuda-graph-max-bs N -> --vllm-max-cudagraph-capture-size N --sglang-cuda-graph-bs a b c -> --vllm-cudagraph-capture-sizes a b c --sglang-ep-size N -> --vllm-enable-expert-parallel --sglang-speculative-* (eagle) -> --vllm-speculative-config '{"method":"eagle","num_speculative_tokens":K}' Also: - glm4.7 pd: fix --vllm-max-num-seqs (was 8, taken from cuda-graph-max-bs; --sglang-max-running-requests was 16) and split out cuda-graph capture. - fix sglang→rollout mis-renames in temp-file prefixes (→ vllm_*). - test_vllm_config: rename test_update_weights_default_true → test_update_weights_defaults_to_none (it asserts `is None`). Dropped sglang flags with no vLLM equivalent (enable-dp-lm-head, moe-dense-tp-size, watchdog-timeout, mamba-scheduler-strategy, disaggregation-transfer-backend, enable-metrics) stay dropped; PD KV-transfer is driven by --prefill-num-servers + the --vllm-config prefill/decode topology. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(plugin_contracts): migrate from sglang_rollout to vllm_rollout The three plugin-contract tests still imported slime.rollout.sglang_rollout and called install_stubs(with_sglang_router=True), but _shared.install_stubs already dropped that parameter — so all three failed at collection (TypeError: unexpected keyword 'with_sglang_router'). Complete the migration: - install_stubs(with_sglang_router=True, ...) -> install_stubs(...) - import generate_and_rm / generate_rollout from slime.rollout.vllm_rollout - default rollout/eval path string -> slime.rollout.vllm_rollout.generate_rollout (matches runtime default at slime/utils/arguments.py:233) - FakeGenerateState: sglang_enable_deterministic_inference -> vllm_enable_deterministic_inference, with group_sampling_seeds defaulting to None and gated on the flag (mirrors the already-migrated tests/unit/rollout/test_vllm_rollout.py). All 34 plugin-contract cases pass (were 3 collection errors before). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(update_weight_from_tensor): drop stale slime…megatron_utils.sglang mock The test pre-registered a sys.modules mock for slime.backends.megatron_utils.sglang (monkey_patch_torch_reductions), left over from when update_weight_from_tensor imported it. The module under test no longer imports that module (its real deps are get_gloo_group / HfWeightIteratorBase / update_weight_from_distributed), so the mock is dead. Removing it makes tests/ and .github/ fully sglang-free. Test still passes (7/7). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * [Clean] Remove SGLang runtime code Rebuilt against current main so the PR contains only the SGLang runtime removal -- the docs / tests-ci / examples / scripts / docker portions were split into separate PRs that have since merged. - Delete dead SGLang server/runtime code: sglang_utils/{arguments,sglang_engine}.py, rollout/sglang_rollout.py, the megatron_utils/sglang.py re-export shim, and all docker/**/sglang.patch files. - Rename the rollout config module sglang_utils/sglang_config.py -> vllm_utils/vllm_config.py (SglangConfig -> VllmConfig, _resolve_sglang_config -> _resolve_vllm_config, --sglang-config -> --vllm-config); inline the GPU_MEMORY_TYPE_* constants in rollout.py. - Add megatron_utils/fp8_helpers.py for the UE8M0 fp8 helpers formerly re-exported through the sglang shim; repoint quantizer_fp8 to it. - Swap sglang_router -> vllm_router in http_utils/wandb_utils; drop the dead sglang-router dependency from requirements.txt. - Finish the SGLang->vLLM rename in the runtime so it is internally consistent and matches the tests landing in the tests/CI PR: * router args --router-* -> --vllm-router-* (vllm_router_ip/port/timeout); * get_model_url reads vllm_model_routers (aligning with rollout.py); * --opd-type sglang -> vllm; engine_overrides rename; * sglang_enable_deterministic_inference -> vllm_enable_deterministic_inference, wired to a real --vllm-enable-deterministic-inference flag (exports VLLM_BATCH_INVARIANT=1); * consistent_hash session-id routing uses vllm-router's x-session-id header; * drop dead trace helper build_sglang_meta_trace_attrs; de-SGLang comments/docstrings. - Rename test_sglang_config.py -> test_vllm_config.py and de-SGLang the plugin-contract tests. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Address review: finish de-SGLang + fold OPD/router-policy into runtime - naming: replace residual generic "rollout engine"/"engine"/"comm" wording with concrete vLLM (engine_overrides -> vllm_overrides; arguments help text; http_utils comments; rollout.py "inference workers"). sglang->vllm is correct, sglang->generic is not. - megatron_to_hf: drop the q_a_proj/kv_a_proj_with_mqa pairing + _cached_tensors global. That was sglang-only: sglang's loader torch.cat's both shards within a single load_weights call (needs them co-bucketed), whereas vLLM loads each shard independently via stacked_params_mapping into fused_qkv_a_proj. Also fix the misleading "merge into single fused name" comment. - docker/Dockerfile: remove now-dead sglang/sglang-router --no-deps stubs + the build-time `import sglang` smoke check (slime no longer imports sglang_router). - OPD: migrate on_policy_distillation.py teacher logprobs to vLLM /v1/completions (prompt_logprobs) instead of sglang return_logprob / meta_info.input_token_logprobs. - routing replay: register --vllm-router-policy (dest=router_policy) so the consistent_hash x-session-id session-affinity path is actually wired (was dead). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Review follow-ups: mirror slime vllm_config parsing + restore vLLM process cleanup - vllm_config.from_yaml: drop the needless `models_raw` intermediate and iterate `data["vllm"]` directly, restoring the "Accept both server_groups / legacy engine_groups" comment -- mirrors slime's sglang_config.from_yaml line-for-line. - command_utils.execute_train: re-add a process kill for leftover rollout engines as `pkill -9 -f "vllm serve"` (the old `pkill -9 sglang` was dropped with no vLLM equivalent), so stale engines don't hold GPUs/ports across runs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Drop test changes from runtime PR; tests live in the tests+CI PR (#40) The plugin_contracts tests and the test_sglang_config -> test_vllm_config rename are coupled to the test/CI rename effort and are owned by #40. Restore them to main here so #18 is purely the SGLang runtime removal. #18 merges first; #40 rebases and re-lands the vLLM test versions. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fp8_helpers: copy SGLang verbatim (fix import crash); pin vLLM deep_gemm env - fp8_helpers.py: replace the bespoke rewrite with SGLang's exact implementations of quant_weight_ue8m0 / transform_scale_ue8m0 and their DeepGEMM helpers (per_block_cast_to_fp8, ceil_to_ue8m0, ceil_div, ceil_align, the torch-impl packer). deep_gemm is imported lazily inside the functions (as SGLang does), so module import no longer requires deep_gemm. This fixes the module-level `NameError: _get_tma_aligned_size` that crashed `import megatron_to_hf` on any deep_gemm image, and drops the invented sf-stride fixup block that was not in upstream. Only should_deepgemm_weight_requant_ue8m0 stays vLLM-adapted (is_deep_gemm_e8m0_used) since SGLang's reads SGLang-internal deep_gemm_wrapper. - vllm_engine.launch_server_process: set VLLM_USE_DEEP_GEMM=1 + VLLM_DEEP_GEMM_WARMUP=relax explicitly (setdefault) alongside VLLM_BATCH_INVARIANT, replacing SGLang's removed deep_gemm precompile/warmup envs. All vLLM engine env now lives in the subprocess env builder (single source of truth). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fp8_helpers: revert to vLLM impl + fix the import NameError; add vLLM trace attrs - fp8_helpers.py: keep the vLLM-based implementation (uses vllm.utils.deep_gemm, consistent with the vLLM runtime) rather than the SGLang verbatim copy. Fix the module-level crash: the `try` block referenced `_get_tma_aligned_size` before it was bound (the "pre-imported with fallback" import was never written), which raised NameError whenever deep_gemm imported successfully -- and NameError is not caught by `except ImportError`, so `import megatron_to_hf` crashed on any deep_gemm image. Replace the bogus self-assignment with the real import: `from vllm.utils.deep_gemm import get_tma_aligned_size as _get_tma_aligned_size`. - trace_utils/vllm_rollout: add build_vllm_meta_trace_attrs and attach finish_reason + token usage to the vllm_inference_generate span (mirrors SGLang's build_sglang_meta_trace_attrs; vLLM responses lack the pd_* timing, which lives in vLLM's own OTLP traces). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(opd): score teacher via /inference/v1/generate with prompt_logprobs Move the vllm OPD teacher path off the OpenAI /v1/completions endpoint onto vime's native /inference/v1/generate (the same endpoint the rollout engines use), and fix three latent issues: 1. model field: /inference/v1/generate takes `model` as OPTIONAL. Stop defaulting to args.hf_checkpoint (the *student* name, which mis-names a teacher!=student server). Add --opd-teacher-model; send `model` only when set, otherwise omit it (single-model teacher servers use their loaded model). 2. multimodal: the old code sent image_data to a token-only endpoint, which is invalid. Raise NotImplementedError until the /v1/chat/completions/render -> /inference/v1/generate flow is wired (mirrors slime.rollout.vllm_rollout.generate). 3. logprob robustness: read top-level GenerateResponse.prompt_logprobs, assert it is present and length-aligned with token_ids, assert the per-sample tensor covers response_length, and raise (not silently return 0.0) on a missing token logprob. vLLM always includes the actual prompt token in prompt_logprobs, so a miss is a real error. Alignment is unchanged (plp[i] <-> tokens[i], skip pos 0, take [-response_length:]). Follow-up (separate, in the tests PR): the OPD e2e test must launch a teacher that exposes /inference/v1/generate and point --rm-url at it. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(clean-sglang): purge SGLang from tools, train scripts, and build infra tools/: drop dead `args.sglang_enable_ep_moe` shim (read nowhere); reword profile/replay helpers to vLLM and map analyzer hints to vLLM flags (--enforce-eager, --gpu-memory-utilization). train{,_async}.py: comments SGLang -> vLLM. build infra: remove build_conda.sh (SGLang-only conda path); drop the GB300 sgl-kernel install from the Dockerfile; delete docker/npu_patch/ wholesale. docker base image: bump to vLLM v0.22.0. justfile ARM recipes now pin the real multi-arch vLLM base images instead of the dead SGLANG_IMAGE_TAG/ ENABLE_SGLANG_PATCH build-args -- cu129-arm64 -> v0.22.0-cu129-ubuntu2404 (CUDA 12.9), cu13-arm64 -> v0.22.0-ubuntu2404 (the default-CUDA tag is already CUDA 13.0) + ENABLE_CUDA_13=1. vLLM tags are multi-arch manifests, so docker selects the arm64 image automatically on an ARM host. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * ci(clean-sglang): drop conda-build workflow (ran deleted build_conda.sh on SGLang image) The single build-conda job ran `bash build_conda.sh` (removed in the previous commit) inside an lmsysorg/sglang container. With the SGLang-only conda path gone, the whole workflow is dead. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(clean-sglang): fix stale SGLang refs in docs/skills; tidy comments docs/conf.py: point the "edit on GitHub" links at vllm-project/vime instead of the inherited sgl-project.github.io repo. .claude/skills/*: update the dead `slime/rollout/sglang_rollout.py` references to `vllm_rollout.py` (the real default is slime.rollout.vllm_rollout.generate_rollout). justfile: drop the redundant BASE_IMAGE override on release-cu129-arm64 (it equalled the Dockerfile default; the multi-arch manifest already resolves arm64). train{,_async}.py: drop stray "the" in the W&B comment. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(tests): use method=mtp (not eagle) in vllm speculative config The migrated speculative configs pass no draft `model`, so method=eagle raises "num_speculative_tokens was provided but without speculative model" in vLLM's SpeculativeConfig. These models carry embedded MTP layers, so method=mtp is correct and unblocks the mimo MTP-only-grad test (#19). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(cleanup): target renamed vLLM subprocesses in pkill so VRAM is freed vLLM's set_process_title() renames the VRAM-holding subprocesses (VLLM::EngineCore, VLLM::Worker_TP*, vllm::router), so their cmdline no longer contains "vllm serve". The previous `pkill -9 -f "vllm serve"` matched only the launcher and left engine/worker children holding GPU memory, leaking it into the next run — masked only by the indiscriminate `pkill -9 python`, which is unsafe on colocate/shared nodes. Match both the launcher and the renamed children with `pkill -9 -f '[v]llm serve|VLL[M]::'`; the [v]/[M] bracket trick keeps the pattern from matching pkill's own cmdline. This makes the broad python kill unnecessary, so its already-commented-out lines are removed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(speculative): use method=mtp (not eagle) for embedded-MTP models vLLM's SpeculativeConfig requires an explicit draft `model` for method=eagle; with only num_speculative_tokens set it raises "num_speculative_tokens was provided but without speculative model". The migrated configs in scripts/examples/docs pass no model, so they must use method=mtp, which reuses the target checkpoint's embedded MTP layer (DeepSeek-R1, GLM-4.x-MoE, MiMo, Qwen3-Next/3.5). The two docs examples that pass an explicit "model" are genuine eagle usage and are left unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(vllm): launch each rollout engine with its ServerGroup's per-group TP launch_server_process / _init_normal derived tensor-parallel size and CUDA_VISIBLE_DEVICES from the global --rollout-num-gpus-per-engine, ignoring the per-engine num_gpus_per_engine already carried on the VLLMEngine actor. A ServerGroup configured with num_gpus_per_engine greater than the global flag (e.g. tp=2) therefore launched as tp=1, while the NCCL weight-sync rendezvous sized world_size from engine_gpu_counts (the per-group value). The two disagreed: the trainer waited for a rank the under-sized engine never started, so init_weight_transfer_engine hung for 300s ("3/4 clients joined") and the job failed. Honor the per-engine num_gpus_per_engine at launch, falling back to the global flag when unset (matches the SGLang path and PR #66's _compute_server_args). Verified on H200: tests/test_qwen2.5_0.5B_vllm_config_distributed now launches engine0 tp=2 / engine1 tp=1, update_weights completes in 1.1s (was a 301s timeout), and rollout+eval proceed. AI assistance (Claude Code) was used for this change. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(ckpt): add --dist-ckpt-optim-fully-reshardable for PAO+offload save/load test_qwen3_4B_ckpt.py uses precision-aware optimizer + cpu-offload (HybridDeviceOptimizer). Under the default dp_reshardable (bucket-centric) optimizer sharding, save/load produce unequal-length param_state lists, so dist-ckpt load fails with "Cannot merge two lists with different lengths (81 and 79)". fully_reshardable is model-centric and immune to bucket-layout changes. Verified on the r3 image (Megatron-LM 0.16.0rc0 @ 1dcf0da): save+load both succeed, and source review confirms master_param / step / HybridDeviceOptimizer sync are handled on this path. This is the flag described in PR #50 that was never actually merged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * refactor(router-args): hybrid naming — vllm_ for ip/port, bare router_ for timeout vllm-router's RouterArgs.from_cli_args only supports prefix "" or "router_" (never "vllm_router_"), and excludes host/port from its CLI via exclude_host_port=True. So: - --vllm-router-ip / --vllm-router-port keep the vllm_ prefix: RouterArgs does not own these CLI flags, vime does (populated via _start_router's manual router_args.host/port assignment), so the vllm_ prefix is free and marks them as vime-owned endpoint config. - --router-request-timeout-secs goes bare (dest router_request_timeout_secs): it is a genuine RouterArgs field, so it shares the --router-* namespace with policy / cache_threshold / retries / … and flows through from_cli_args like the other knobs. - --vllm-router-policy keeps dest=router_policy (unchanged). Also fixes conftest fixture to seed vllm_router_ip/port (was bare router_ip/port, which never matched the vllm_engine reader) and updates README/README_zh prose. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: Canlin Guo <canlinguosdu@gmail.com>
Summary
Split from PR #18 (
gcl/clean-sglang). One of 4 PRs splitting the original PR #18 by content area: docs / examples / tests+CI / core runtime. See pr18-split-plan.md on PR #18.53 files. Pure documentation: no runtime code change.
What this PR does
docs/en/+docs/zh/(49 files):sglang-config.md→vllm-config.md(1 each side)--sglang-*CLI flag examples to their--vllm-*equivalentsblog.vllm.ai/2025/10/26/distributed-batch-invariance.htmlvllm.ai/blog/2025-11-10-bitwise-consistent-train-inferencedocs.vllm.ai/.../features/spec_decode.html(301 → 404)docs.vllm.ai/.../features/speculative_decoding/github.com/vllm-project/vllm-router(404; repo does not exist)github.com/vllm-project/router(the actual router repo)docs/_static/image/(3 files): renamesglang_config.png→vllm_config.png; refresharch.pngto the vllm-router architecture.docs/{en,zh}/advanced/reproducibility.md: The en and zh files had their content swapped (en had Chinese body, zh had English body). Restored the correct mapping.docs/{en,zh}/blogs/introducing_slime.md: Remove the "Special thanks to the AMD GenAI - Foundation Model Team for Day-1 AMD hardware support" line. With AMD support removed in docker/scripts/runtime: remove AMD / ROCm platform support #36, the acknowledgement implied ongoing support that no longer exists.README.md/README_zh.md:--router-*for vllm-router native flags vs--vllm-router-*for vime-side orchestration knobs)docs/zh/advanced/slime_vllm_backend_design_v1.md(deleted in commitfe700db)Independent of other splits
This PR can land at any time without coordination. See
pr18-split-plan.md§"合并顺序无强约束".Test plan
git grep -E 'spec_decode\.html|vllm-project/vllm-router|distributed-batch-invariance' -- docs/returns no matchesgit grep 'AMD GenAI' -- docs/returns no matchescd docs && make html— no broken cross-references🤖 Generated with Claude Code