Skip to content

[WIP][Feature] Optional vLLM router for internal rollout gateway - #1

Closed
CalvinXKY wants to merge 1 commit into
mainfrom
xky-dev
Closed

[WIP][Feature] Optional vLLM router for internal rollout gateway#1
CalvinXKY wants to merge 1 commit into
mainfrom
xky-dev

Conversation

@CalvinXKY

Copy link
Copy Markdown
Collaborator

Purpose

Add support for spawning vllm-router as the HTTP rollout gateway in addition to the existing sglang-router (sgl-model-gateway) path.

  • New flag --router-impl {sglang,vllm} (args.router_impl): choose which gateway binary Slime starts when no external router is configured (--sglang-router-ip unset).
  • New module slime.utils.router: register shared --router-* / --worker-urls CLI, build RouterArgs, assert the selected package is installed, map worker control-plane to the modern /workers API when router_impl=vllm, and sanitize RouterArgs for vLLM when both wheels are installed (e.g. sglang’s max_concurrent_requests=-1 is invalid for vLLM’s unsigned Rust fields).
  • slime/utils/http_utils.py: run_router dispatches to vllm_router.launch_router or sglang_router.launch_router based on impl.
  • requirements.txt: add vllm-router>=0.1.14 alongside sglang-router.
    Default behavior remains sglang; external routers are unchanged.

Test Plan

  • Install deps: pip install -r requirements.txt
  • Train / rollout with --router-impl vllm (internal router) and confirm log line: Launch HTTP router (impl=vllm) ...
  • Smoke with --router-impl sglang (or omit flag) to confirm no regression

running:

export PYTHONPATH=/root/Megatron-LM
SCRIPT_DIR="/data/nfs/kaiyuan/RL/vime/scripts"
source "${SCRIPT_DIR}/models/qwen3-0.6B.sh"

python train.py \
  ${MODEL_ARGS[@]} \
  \
  --hf-checkpoint /data/nfs/kaiyuan/models/Qwen3-0.6B \
  --ref-load /data/nfs/kaiyuan/models/Qwen3-0.6B_torch_dist \
  --load /work/Qwen3-0.6B_slime_ckpt \
  --save /work/Qwen3-0.6B_slime_ckpt \
  --save-interval 10 \
  --train-memory-margin-bytes 2147483648 \
  \
  --prompt-data /data/nfs/kaiyuan/datasets/dapo-math-17k/dapo-math-17k.jsonl \
  --input-key prompt \
  --label-key label \
  --apply-chat-template \
  --rollout-shuffle \
  \
  --rm-type deepscaler \
  --num-rollout 100 \
  --rollout-batch-size 4 \
  --n-samples-per-prompt 2 \
  --num-steps-per-rollout 1 \
  --global-batch-size 8 \
  \
  --actor-num-nodes 1 \
  --actor-num-gpus-per-node 8 \
  --rollout-num-gpus 8 \
  --rollout-num-gpus-per-engine 2 \
  --rollout-max-response-len 1024 \
  --sglang-mem-fraction-static 0.8 \
  \
  --attention-dropout 0.0 \
  --hidden-dropout 0.0 \
  --accumulate-allreduce-grads-in-fp32 \
  --attention-softmax-in-fp32 \
  --attention-backend flash \
  --optimizer adam \
  --lr 1e-6 \
  --lr-decay-style constant \
  --weight-decay 0.1 \
  --adam-beta1 0.9 \
  --adam-beta2 0.98 \
  --router-impl vllm \
  \
  --colocate

Test Result

image

@CalvinXKY CalvinXKY changed the title [Feature] Optional vLLM router for internal rollout gateway [WIP][Feature] Optional vLLM router for internal rollout gateway May 14, 2026
@aoshen02 aoshen02 closed this May 16, 2026
@aoshen02 aoshen02 mentioned this pull request May 18, 2026
14 tasks
aoshen02 added a commit that referenced this pull request May 28, 2026
The 786-line tests/test_update_weight_from_tensor.py is a stale rebase
leftover from the original PR #18 branch — it predates the IPC test
file PR #22 landed at the canonical unit-test path
(tests/unit/backends/megatron_utils/update_weight/test_update_weight_from_tensor.py)
and predates PR #48's single-RPC weight-version contract.

Comparing the two:
* Both stub sys.modules / torch.distributed at module import time, so
  having two files compounds the test-isolation issue Gemini raised
  (PR #40 comment #1).
* Coverage overlaps materially (e.g. test_ipc_init_called_on_first_update_only
  ≈ test_ipc_init_runs_once — same invariant, different wording).
* The nested file is up-to-date with PR #48's RPC contract
  (update_weights_from_tensor.remote(**fields, weight_version=...));
  the top-level file still uses the pre-#48 lifecycle shape and does
  not exercise the coordinator slot fields.
* The nested path matches repo convention: tests/unit/ for mock-only
  unit tests, tests/ top level for e2e scripts.

Closes Gemini comment #1 on PR #40. Gemini comment #2 (the same stub
pattern in the surviving nested file) is a pre-existing issue from
PR #22 / #48 and out of scope for this rename PR — to be addressed
in a follow-up that converts _install_stubs() to an autouse
module-scoped fixture with save/restore.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
CalvinXKY pushed a commit that referenced this pull request May 30, 2026
* tests + CI: complete sglang→vllm rename across tests/ and .github/

Split from PR #18 (gcl/clean-sglang). One of 4 PRs splitting the
original PR #18 by content area: docs (#38) / examples (#39) /
**tests+CI** / core runtime. 42 files / +~750 / -~700.

These are bundled in a single PR because the CI workflows reference
test file names by string — splitting them would create a window where
either tests are renamed but CI still points at the old names, or vice
versa, breaking CI mid-roll.

What this PR does:

(A) tests/ (38 files):

- Mechanical CLI-flag rename: --sglang-* → --vllm-* equivalents in all
  test scripts (matches the table now used in scripts/ and examples/).
- Variable rename: SGLANG_ARGS → VLLM_ARGS where present.
- 4 file renames (R086-R091, all >85% similarity):
    test_qwen2.5_0.5B_opd_sglang.py        → test_qwen2.5_0.5B_opd_vllm.py
    test_qwen2.5_0.5B_sglang_config.py     → test_qwen2.5_0.5B_vllm_config.py
    test_qwen2.5_0.5B_sglang_config_distributed.py
                                           → test_qwen2.5_0.5B_vllm_config_distributed.py
    test_sglang_config_mixed_offload.py    → test_vllm_config_mixed_offload.py
    test_sglang_config_mixed_offload_ft.py → test_vllm_config_mixed_offload_ft.py
    tests/utils/test_sglang_config.py      → tests/utils/test_vllm_config.py
- 2 new tests for the IPC weight-transfer path landed in PR #18:
    tests/test_update_weight_from_tensor.py
    tests/unit/backends/megatron_utils/update_weight/test_update_weight_from_tensor.py
  (These are PR #22 / colocate-IPC test coverage; the production code
  the slim PR #18 ships will rely on the same code from PR #22.)

(B) .github/ (4 files):

- workflows/conda-ci.yml: container image lmsysorg/sglang → vime
  (inferactinc/public:vime-vllm-cu129-latest).
- workflows/pr-test.yml + pr-test.yml.j2 (template):
    * Container images (slimerl/slime[-test]:latest → vime image) on
      every job that ran on the sglang-era base.
    * e2e-test-sglang-config job → e2e-test-vllm-config job (renamed
      label `run-ci-sglang-config` → `run-ci-vllm-config`; matrix
      `test_file` entries updated to point at the renamed test files
      in (A)).
    * e2e-test-megatron + e2e-test-image matrices: `_opd_sglang.py`
      entries → `_opd_vllm.py`.
- ISSUE_TEMPLATE/bug_report.yml: drop the "SGLang version (if
  relevant):" environment field, add "vLLM version:" and
  "vllm-router version:" lines. (PR #36 already changed
  "CUDA/ROCm version" → "CUDA version" earlier; that change is
  preserved.)

Sgl residue intentionally kept (4 hits — all anti-regression
assertions that prove sglang code paths are gone, not residual
references to bring back):

- tests/test_update_weight_from_tensor.py:753 — comment "The vLLM IPC
  implementation must NOT contain sglang-style Gloo gather code".
- tests/unit/backends/vllm_utils/test_arguments.py:233-237 — three
  assertions that --sglang-router-ip, --sglang-router-port, and
  sglang_router_ip are NOT present in the argument parser.

Tests + CI must land together; splitting them risks a window where
the CI matrix references test files by names that don't exist yet
(or no longer exist). After this lands, the test_file string in CI
matches the test files on disk.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Canlin Guo <canlinguosdu@gmail.com>

* on_policy_distillation: port from SGLang to vLLM /v1/completions

Follow-up on the test rename in this PR:
test_qwen2.5_0.5B_opd_sglang.py → test_qwen2.5_0.5B_opd_vllm.py.

The test only spawns a vLLM teacher and exercises the OPD pipeline; the
real broken piece was slime/rollout/on_policy_distillation.py, which
PR #18 left in SGLang request/response shape:

  request fields:
    "max_new_tokens": 0          (vLLM: "max_tokens")
    "return_logprob": True       (sglang-only)
    "logprob_start_len": 0       (sglang-only)
  response parsing:
    reward["meta_info"]["input_token_logprobs"]   (sglang shape)

vLLM 0.21 supports the same workflow natively via `prompt_logprobs`:

  request to POST /v1/completions:
    {
        "model": <teacher>,
        "prompt_token_ids": sample.tokens,
        "max_tokens": 1,
        "temperature": 0,
        "prompt_logprobs": 1,
        "logprobs": 0,
        "skip_special_tokens": False,
    }
  response:
    response["choices"][0]["prompt_logprobs"]   # list[dict[int, Logprob] | None]
      where Logprob is {"logprob": float, "rank": int, "decoded_token": str}

References checked against vllm source:
  - reference/vllm/vllm/entrypoints/openai/completion/protocol.py:91
    (request: prompt_logprobs: int | None)
  - reference/vllm/vllm/entrypoints/openai/completion/protocol.py:487
    (response: prompt_logprobs: list[dict[int, Logprob] | None] | None)
  - reference/vllm/vllm/logprobs.py:13
    (Logprob dataclass: logprob/rank/decoded_token)

Implementation notes:

1. JSON serializes int dict keys as strings, so `_logprob_for_token`
   tries both `pos_entry.get(token_id)` and `pos_entry.get(str(token_id))`.

2. `pos_entry` is `None` at position 0 (no prior context) — handled
   explicitly. We also gracefully degrade if a token at position `i` is
   not in the top-1 logprob dict (falls back to 0.0, same as the prior
   sglang code would do).

3. The Logprob dataclass `decoded_token` field is unused; we only read
   `.logprob`. Both dict and `Logprob` shapes are accepted in case the
   server uses a flatter serialization toggle.

4. `args.opd_teacher_model` is the new model-name arg; falls back to
   `args.hf_checkpoint` if not set, mirroring how vime's other rollout
   paths derive the model name.

Smoke-tested `_logprob_for_token` locally:
  - None entry → 0.0
  - int key + dict value → logprob
  - str key (JSON shape) → logprob
  - missing token → 0.0
  - flattened float value → float

Also drops 3 lines from
tests/unit/backends/vllm_utils/test_arguments.py: the
`--sglang-router-ip`/`--sglang-router-port`/`sglang_router_ip` anti-
regression assertions. Once the slim PR #18 lands and sglang is gone
from the runtime, those assertions are vacuous; treating sglang as
non-existent per the project policy.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Canlin Guo <canlinguosdu@gmail.com>

* tests: drop duplicate smoke test updates from PR40

* test(update_weight_from_tensor): drop stale _apply_monkey_patch_torch_reductions patch

The inner ``with patch(f"{MODULE_PATH}._apply_monkey_patch_torch_reductions"):``
context in _run_update suppressed a helper call that PR #48 has since deleted
from update_weight_from_tensor.py (commit 39bf899 on aoshen/align-ipc-rpc-with-slime).
After that PR lands the patched attribute won't exist and this line raises
AttributeError. Remove it now so the test survives PR #48 merge.

The ``sglang_mod.monkey_patch_torch_reductions = MagicMock()`` stub on the
fake sglang module is intentionally kept: on this branch the production code
still imports it via ``from ..sglang import monkey_patch_torch_reductions``
(both update_weight_from_tensor._apply_monkey_patch_torch_reductions on
PR #40's view of main, and hf_weight_iterator_direct.py at module level).
Removing the stub here would break the test on PR #40 alone; it can be
dropped in a follow-up once PR #48 finishes removing every import site.

Tests: ``tests/unit/backends/megatron_utils/update_weight/test_update_weight_from_tensor.py``
all 6 pass with this change applied to gcl/pr18-tests-ci HEAD.

* tests: drop duplicate top-level test_update_weight_from_tensor.py

The 786-line tests/test_update_weight_from_tensor.py is a stale rebase
leftover from the original PR #18 branch — it predates the IPC test
file PR #22 landed at the canonical unit-test path
(tests/unit/backends/megatron_utils/update_weight/test_update_weight_from_tensor.py)
and predates PR #48's single-RPC weight-version contract.

Comparing the two:
* Both stub sys.modules / torch.distributed at module import time, so
  having two files compounds the test-isolation issue Gemini raised
  (PR #40 comment #1).
* Coverage overlaps materially (e.g. test_ipc_init_called_on_first_update_only
  ≈ test_ipc_init_runs_once — same invariant, different wording).
* The nested file is up-to-date with PR #48's RPC contract
  (update_weights_from_tensor.remote(**fields, weight_version=...));
  the top-level file still uses the pre-#48 lifecycle shape and does
  not exercise the coordinator slot fields.
* The nested path matches repo convention: tests/unit/ for mock-only
  unit tests, tests/ top level for e2e scripts.

Closes Gemini comment #1 on PR #40. Gemini comment #2 (the same stub
pattern in the surviving nested file) is a pre-existing issue from
PR #22 / #48 and out of scope for this rename PR — to be addressed
in a follow-up that converts _install_stubs() to an autouse
module-scoped fixture with save/restore.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(vllm_config): use real get_model_url default endpoint /inference/v1/generate

get_model_url defaults to /inference/v1/generate (PR #18), not /v1/completions.
Aligns this test with PR #18's test_vllm_config.py so the two PRs no longer
conflict on this file and the assertion matches the actual runtime default.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Drop on_policy_distillation.py from tests+CI PR (now owned by runtime PR #18)

The OPD vLLM /v1/completions migration is a runtime change; it was folded into
the core-runtime PR (#18). Restore this file to main here so the two PRs no longer
overlap on it. #18 merges first, so this lands via #18.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Drop unit-test files now owned by runtime PR #18

test_vllm_config.py + the plugin_contracts tests are coupled to #18's runtime
rename (they import vllm_config / vllm_rollout, which #18 creates). They live in
#18; remove them here so the two PRs don't overlap. #18 merges first.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: restore vLLM rollout args dropped during sglang→vllm rename

The mechanical sglang→vllm rename dropped several rollout knobs instead of
mapping them to their vLLM equivalents, weakening CI coverage (cuda-graph
capture caps, speculative decoding, expert parallel). Restore them using the
mapping established by the converted production scripts on main
(run-glm4.7-30B-A3B.sh / run-glm5-744B-A40B.sh), verified against vLLM
AsyncEngineArgs:

  --sglang-cuda-graph-max-bs N            -> --vllm-max-cudagraph-capture-size N
  --sglang-cuda-graph-bs a b c            -> --vllm-cudagraph-capture-sizes a b c
  --sglang-ep-size N                      -> --vllm-enable-expert-parallel
  --sglang-speculative-* (eagle)          -> --vllm-speculative-config '{"method":"eagle","num_speculative_tokens":K}'

Also:
- glm4.7 pd: fix --vllm-max-num-seqs (was 8, taken from cuda-graph-max-bs;
  --sglang-max-running-requests was 16) and split out cuda-graph capture.
- fix sglang→rollout mis-renames in temp-file prefixes (→ vllm_*).
- test_vllm_config: rename test_update_weights_default_true →
  test_update_weights_defaults_to_none (it asserts `is None`).

Dropped sglang flags with no vLLM equivalent (enable-dp-lm-head,
moe-dense-tp-size, watchdog-timeout, mamba-scheduler-strategy,
disaggregation-transfer-backend, enable-metrics) stay dropped; PD KV-transfer
is driven by --prefill-num-servers + the --vllm-config prefill/decode topology.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(plugin_contracts): migrate from sglang_rollout to vllm_rollout

The three plugin-contract tests still imported slime.rollout.sglang_rollout
and called install_stubs(with_sglang_router=True), but _shared.install_stubs
already dropped that parameter — so all three failed at collection
(TypeError: unexpected keyword 'with_sglang_router'). Complete the migration:

- install_stubs(with_sglang_router=True, ...) -> install_stubs(...)
- import generate_and_rm / generate_rollout from slime.rollout.vllm_rollout
- default rollout/eval path string -> slime.rollout.vllm_rollout.generate_rollout
  (matches runtime default at slime/utils/arguments.py:233)
- FakeGenerateState: sglang_enable_deterministic_inference ->
  vllm_enable_deterministic_inference, with group_sampling_seeds defaulting to
  None and gated on the flag (mirrors the already-migrated
  tests/unit/rollout/test_vllm_rollout.py).

All 34 plugin-contract cases pass (were 3 collection errors before).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(update_weight_from_tensor): drop stale slime…megatron_utils.sglang mock

The test pre-registered a sys.modules mock for
slime.backends.megatron_utils.sglang (monkey_patch_torch_reductions), left over
from when update_weight_from_tensor imported it. The module under test no longer
imports that module (its real deps are get_gloo_group / HfWeightIteratorBase /
update_weight_from_distributed), so the mock is dead. Removing it makes tests/
and .github/ fully sglang-free. Test still passes (7/7).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [Clean] Remove SGLang runtime code

Rebuilt against current main so the PR contains only the SGLang runtime
removal -- the docs / tests-ci / examples / scripts / docker portions were
split into separate PRs that have since merged.

- Delete dead SGLang server/runtime code: sglang_utils/{arguments,sglang_engine}.py,
  rollout/sglang_rollout.py, the megatron_utils/sglang.py re-export shim, and all
  docker/**/sglang.patch files.
- Rename the rollout config module sglang_utils/sglang_config.py ->
  vllm_utils/vllm_config.py (SglangConfig -> VllmConfig, _resolve_sglang_config ->
  _resolve_vllm_config, --sglang-config -> --vllm-config); inline the
  GPU_MEMORY_TYPE_* constants in rollout.py.
- Add megatron_utils/fp8_helpers.py for the UE8M0 fp8 helpers formerly re-exported
  through the sglang shim; repoint quantizer_fp8 to it.
- Swap sglang_router -> vllm_router in http_utils/wandb_utils; drop the dead
  sglang-router dependency from requirements.txt.
- Finish the SGLang->vLLM rename in the runtime so it is internally consistent and
  matches the tests landing in the tests/CI PR:
  * router args --router-* -> --vllm-router-* (vllm_router_ip/port/timeout);
  * get_model_url reads vllm_model_routers (aligning with rollout.py);
  * --opd-type sglang -> vllm; engine_overrides rename;
  * sglang_enable_deterministic_inference -> vllm_enable_deterministic_inference,
    wired to a real --vllm-enable-deterministic-inference flag (exports
    VLLM_BATCH_INVARIANT=1);
  * consistent_hash session-id routing uses vllm-router's x-session-id header;
  * drop dead trace helper build_sglang_meta_trace_attrs; de-SGLang comments/docstrings.
- Rename test_sglang_config.py -> test_vllm_config.py and de-SGLang the
  plugin-contract tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Address review: finish de-SGLang + fold OPD/router-policy into runtime

- naming: replace residual generic "rollout engine"/"engine"/"comm" wording
  with concrete vLLM (engine_overrides -> vllm_overrides; arguments help text;
  http_utils comments; rollout.py "inference workers"). sglang->vllm is correct,
  sglang->generic is not.
- megatron_to_hf: drop the q_a_proj/kv_a_proj_with_mqa pairing + _cached_tensors
  global. That was sglang-only: sglang's loader torch.cat's both shards within a
  single load_weights call (needs them co-bucketed), whereas vLLM loads each shard
  independently via stacked_params_mapping into fused_qkv_a_proj. Also fix the
  misleading "merge into single fused name" comment.
- docker/Dockerfile: remove now-dead sglang/sglang-router --no-deps stubs + the
  build-time `import sglang` smoke check (slime no longer imports sglang_router).
- OPD: migrate on_policy_distillation.py teacher logprobs to vLLM /v1/completions
  (prompt_logprobs) instead of sglang return_logprob / meta_info.input_token_logprobs.
- routing replay: register --vllm-router-policy (dest=router_policy) so the
  consistent_hash x-session-id session-affinity path is actually wired (was dead).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Review follow-ups: mirror slime vllm_config parsing + restore vLLM process cleanup

- vllm_config.from_yaml: drop the needless `models_raw` intermediate and iterate
  `data["vllm"]` directly, restoring the "Accept both server_groups / legacy
  engine_groups" comment -- mirrors slime's sglang_config.from_yaml line-for-line.
- command_utils.execute_train: re-add a process kill for leftover rollout engines
  as `pkill -9 -f "vllm serve"` (the old `pkill -9 sglang` was dropped with no
  vLLM equivalent), so stale engines don't hold GPUs/ports across runs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Drop test changes from runtime PR; tests live in the tests+CI PR (#40)

The plugin_contracts tests and the test_sglang_config -> test_vllm_config rename
are coupled to the test/CI rename effort and are owned by #40. Restore them to
main here so #18 is purely the SGLang runtime removal. #18 merges first; #40
rebases and re-lands the vLLM test versions.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fp8_helpers: copy SGLang verbatim (fix import crash); pin vLLM deep_gemm env

- fp8_helpers.py: replace the bespoke rewrite with SGLang's exact implementations
  of quant_weight_ue8m0 / transform_scale_ue8m0 and their DeepGEMM helpers
  (per_block_cast_to_fp8, ceil_to_ue8m0, ceil_div, ceil_align, the torch-impl
  packer). deep_gemm is imported lazily inside the functions (as SGLang does), so
  module import no longer requires deep_gemm. This fixes the module-level
  `NameError: _get_tma_aligned_size` that crashed `import megatron_to_hf` on any
  deep_gemm image, and drops the invented sf-stride fixup block that was not in
  upstream. Only should_deepgemm_weight_requant_ue8m0 stays vLLM-adapted
  (is_deep_gemm_e8m0_used) since SGLang's reads SGLang-internal deep_gemm_wrapper.
- vllm_engine.launch_server_process: set VLLM_USE_DEEP_GEMM=1 +
  VLLM_DEEP_GEMM_WARMUP=relax explicitly (setdefault) alongside VLLM_BATCH_INVARIANT,
  replacing SGLang's removed deep_gemm precompile/warmup envs. All vLLM engine env
  now lives in the subprocess env builder (single source of truth).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fp8_helpers: revert to vLLM impl + fix the import NameError; add vLLM trace attrs

- fp8_helpers.py: keep the vLLM-based implementation (uses vllm.utils.deep_gemm,
  consistent with the vLLM runtime) rather than the SGLang verbatim copy. Fix the
  module-level crash: the `try` block referenced `_get_tma_aligned_size` before it
  was bound (the "pre-imported with fallback" import was never written), which
  raised NameError whenever deep_gemm imported successfully -- and NameError is not
  caught by `except ImportError`, so `import megatron_to_hf` crashed on any
  deep_gemm image. Replace the bogus self-assignment with the real import:
  `from vllm.utils.deep_gemm import get_tma_aligned_size as _get_tma_aligned_size`.
- trace_utils/vllm_rollout: add build_vllm_meta_trace_attrs and attach finish_reason
  + token usage to the vllm_inference_generate span (mirrors SGLang's
  build_sglang_meta_trace_attrs; vLLM responses lack the pd_* timing, which lives
  in vLLM's own OTLP traces).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(opd): score teacher via /inference/v1/generate with prompt_logprobs

Move the vllm OPD teacher path off the OpenAI /v1/completions endpoint onto
vime's native /inference/v1/generate (the same endpoint the rollout engines
use), and fix three latent issues:

1. model field: /inference/v1/generate takes `model` as OPTIONAL. Stop
   defaulting to args.hf_checkpoint (the *student* name, which mis-names a
   teacher!=student server). Add --opd-teacher-model; send `model` only when
   set, otherwise omit it (single-model teacher servers use their loaded model).
2. multimodal: the old code sent image_data to a token-only endpoint, which is
   invalid. Raise NotImplementedError until the
   /v1/chat/completions/render -> /inference/v1/generate flow is wired (mirrors
   slime.rollout.vllm_rollout.generate).
3. logprob robustness: read top-level GenerateResponse.prompt_logprobs, assert
   it is present and length-aligned with token_ids, assert the per-sample tensor
   covers response_length, and raise (not silently return 0.0) on a missing
   token logprob. vLLM always includes the actual prompt token in
   prompt_logprobs, so a miss is a real error.

Alignment is unchanged (plp[i] <-> tokens[i], skip pos 0, take [-response_length:]).

Follow-up (separate, in the tests PR): the OPD e2e test must launch a teacher
that exposes /inference/v1/generate and point --rm-url at it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(clean-sglang): purge SGLang from tools, train scripts, and build infra

tools/: drop dead `args.sglang_enable_ep_moe` shim (read nowhere); reword
profile/replay helpers to vLLM and map analyzer hints to vLLM flags
(--enforce-eager, --gpu-memory-utilization). train{,_async}.py: comments
SGLang -> vLLM.

build infra: remove build_conda.sh (SGLang-only conda path); drop the GB300
sgl-kernel install from the Dockerfile; delete docker/npu_patch/ wholesale.

docker base image: bump to vLLM v0.22.0. justfile ARM recipes now pin the real
multi-arch vLLM base images instead of the dead SGLANG_IMAGE_TAG/
ENABLE_SGLANG_PATCH build-args -- cu129-arm64 -> v0.22.0-cu129-ubuntu2404
(CUDA 12.9), cu13-arm64 -> v0.22.0-ubuntu2404 (the default-CUDA tag is already
CUDA 13.0) + ENABLE_CUDA_13=1. vLLM tags are multi-arch manifests, so docker
selects the arm64 image automatically on an ARM host.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(clean-sglang): drop conda-build workflow (ran deleted build_conda.sh on SGLang image)

The single build-conda job ran `bash build_conda.sh` (removed in the previous
commit) inside an lmsysorg/sglang container. With the SGLang-only conda path
gone, the whole workflow is dead.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(clean-sglang): fix stale SGLang refs in docs/skills; tidy comments

docs/conf.py: point the "edit on GitHub" links at vllm-project/vime instead of
the inherited sgl-project.github.io repo. .claude/skills/*: update the dead
`slime/rollout/sglang_rollout.py` references to `vllm_rollout.py` (the real
default is slime.rollout.vllm_rollout.generate_rollout).

justfile: drop the redundant BASE_IMAGE override on release-cu129-arm64 (it
equalled the Dockerfile default; the multi-arch manifest already resolves
arm64). train{,_async}.py: drop stray "the" in the W&B comment.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(tests): use method=mtp (not eagle) in vllm speculative config

The migrated speculative configs pass no draft `model`, so method=eagle
raises "num_speculative_tokens was provided but without speculative model"
in vLLM's SpeculativeConfig. These models carry embedded MTP layers, so
method=mtp is correct and unblocks the mimo MTP-only-grad test (#19).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(cleanup): target renamed vLLM subprocesses in pkill so VRAM is freed

vLLM's set_process_title() renames the VRAM-holding subprocesses
(VLLM::EngineCore, VLLM::Worker_TP*, vllm::router), so their cmdline no
longer contains "vllm serve". The previous `pkill -9 -f "vllm serve"`
matched only the launcher and left engine/worker children holding GPU
memory, leaking it into the next run — masked only by the indiscriminate
`pkill -9 python`, which is unsafe on colocate/shared nodes.

Match both the launcher and the renamed children with
`pkill -9 -f '[v]llm serve|VLL[M]::'`; the [v]/[M] bracket trick keeps the
pattern from matching pkill's own cmdline. This makes the broad python
kill unnecessary, so its already-commented-out lines are removed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(speculative): use method=mtp (not eagle) for embedded-MTP models

vLLM's SpeculativeConfig requires an explicit draft `model` for
method=eagle; with only num_speculative_tokens set it raises
"num_speculative_tokens was provided but without speculative model".
The migrated configs in scripts/examples/docs pass no model, so they must
use method=mtp, which reuses the target checkpoint's embedded MTP layer
(DeepSeek-R1, GLM-4.x-MoE, MiMo, Qwen3-Next/3.5).

The two docs examples that pass an explicit "model" are genuine eagle
usage and are left unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(vllm): launch each rollout engine with its ServerGroup's per-group TP

launch_server_process / _init_normal derived tensor-parallel size and
CUDA_VISIBLE_DEVICES from the global --rollout-num-gpus-per-engine, ignoring
the per-engine num_gpus_per_engine already carried on the VLLMEngine actor.

A ServerGroup configured with num_gpus_per_engine greater than the global
flag (e.g. tp=2) therefore launched as tp=1, while the NCCL weight-sync
rendezvous sized world_size from engine_gpu_counts (the per-group value).
The two disagreed: the trainer waited for a rank the under-sized engine
never started, so init_weight_transfer_engine hung for 300s
("3/4 clients joined") and the job failed.

Honor the per-engine num_gpus_per_engine at launch, falling back to the
global flag when unset (matches the SGLang path and PR #66's
_compute_server_args).

Verified on H200: tests/test_qwen2.5_0.5B_vllm_config_distributed now
launches engine0 tp=2 / engine1 tp=1, update_weights completes in 1.1s
(was a 301s timeout), and rollout+eval proceed.

AI assistance (Claude Code) was used for this change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(ckpt): add --dist-ckpt-optim-fully-reshardable for PAO+offload save/load

test_qwen3_4B_ckpt.py uses precision-aware optimizer + cpu-offload
(HybridDeviceOptimizer). Under the default dp_reshardable (bucket-centric)
optimizer sharding, save/load produce unequal-length param_state lists, so
dist-ckpt load fails with
"Cannot merge two lists with different lengths (81 and 79)".

fully_reshardable is model-centric and immune to bucket-layout changes.
Verified on the r3 image (Megatron-LM 0.16.0rc0 @ 1dcf0da): save+load both
succeed, and source review confirms master_param / step / HybridDeviceOptimizer
sync are handled on this path. This is the flag described in PR #50 that was
never actually merged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(router-args): hybrid naming — vllm_ for ip/port, bare router_ for timeout

vllm-router's RouterArgs.from_cli_args only supports prefix "" or "router_"
(never "vllm_router_"), and excludes host/port from its CLI via
exclude_host_port=True. So:

- --vllm-router-ip / --vllm-router-port keep the vllm_ prefix: RouterArgs does
  not own these CLI flags, vime does (populated via _start_router's manual
  router_args.host/port assignment), so the vllm_ prefix is free and marks them
  as vime-owned endpoint config.
- --router-request-timeout-secs goes bare (dest router_request_timeout_secs): it
  is a genuine RouterArgs field, so it shares the --router-* namespace with
  policy / cache_threshold / retries / … and flows through from_cli_args like
  the other knobs.
- --vllm-router-policy keeps dest=router_policy (unchanged).

Also fixes conftest fixture to seed vllm_router_ip/port (was bare router_ip/port,
which never matched the vllm_engine reader) and updates README/README_zh prose.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Canlin Guo <canlinguosdu@gmail.com>
momo609 pushed a commit that referenced this pull request Jun 8, 2026
* tests + CI: complete sglang→vllm rename across tests/ and .github/

Split from PR #18 (gcl/clean-sglang). One of 4 PRs splitting the
original PR #18 by content area: docs (#38) / examples (#39) /
**tests+CI** / core runtime. 42 files / +~750 / -~700.

These are bundled in a single PR because the CI workflows reference
test file names by string — splitting them would create a window where
either tests are renamed but CI still points at the old names, or vice
versa, breaking CI mid-roll.

What this PR does:

(A) tests/ (38 files):

- Mechanical CLI-flag rename: --sglang-* → --vllm-* equivalents in all
  test scripts (matches the table now used in scripts/ and examples/).
- Variable rename: SGLANG_ARGS → VLLM_ARGS where present.
- 4 file renames (R086-R091, all >85% similarity):
    test_qwen2.5_0.5B_opd_sglang.py        → test_qwen2.5_0.5B_opd_vllm.py
    test_qwen2.5_0.5B_sglang_config.py     → test_qwen2.5_0.5B_vllm_config.py
    test_qwen2.5_0.5B_sglang_config_distributed.py
                                           → test_qwen2.5_0.5B_vllm_config_distributed.py
    test_sglang_config_mixed_offload.py    → test_vllm_config_mixed_offload.py
    test_sglang_config_mixed_offload_ft.py → test_vllm_config_mixed_offload_ft.py
    tests/utils/test_sglang_config.py      → tests/utils/test_vllm_config.py
- 2 new tests for the IPC weight-transfer path landed in PR #18:
    tests/test_update_weight_from_tensor.py
    tests/unit/backends/megatron_utils/update_weight/test_update_weight_from_tensor.py
  (These are PR #22 / colocate-IPC test coverage; the production code
  the slim PR #18 ships will rely on the same code from PR #22.)

(B) .github/ (4 files):

- workflows/conda-ci.yml: container image lmsysorg/sglang → vime
  (inferactinc/public:vime-vllm-cu129-latest).
- workflows/pr-test.yml + pr-test.yml.j2 (template):
    * Container images (slimerl/slime[-test]:latest → vime image) on
      every job that ran on the sglang-era base.
    * e2e-test-sglang-config job → e2e-test-vllm-config job (renamed
      label `run-ci-sglang-config` → `run-ci-vllm-config`; matrix
      `test_file` entries updated to point at the renamed test files
      in (A)).
    * e2e-test-megatron + e2e-test-image matrices: `_opd_sglang.py`
      entries → `_opd_vllm.py`.
- ISSUE_TEMPLATE/bug_report.yml: drop the "SGLang version (if
  relevant):" environment field, add "vLLM version:" and
  "vllm-router version:" lines. (PR #36 already changed
  "CUDA/ROCm version" → "CUDA version" earlier; that change is
  preserved.)

Sgl residue intentionally kept (4 hits — all anti-regression
assertions that prove sglang code paths are gone, not residual
references to bring back):

- tests/test_update_weight_from_tensor.py:753 — comment "The vLLM IPC
  implementation must NOT contain sglang-style Gloo gather code".
- tests/unit/backends/vllm_utils/test_arguments.py:233-237 — three
  assertions that --sglang-router-ip, --sglang-router-port, and
  sglang_router_ip are NOT present in the argument parser.

Tests + CI must land together; splitting them risks a window where
the CI matrix references test files by names that don't exist yet
(or no longer exist). After this lands, the test_file string in CI
matches the test files on disk.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Canlin Guo <canlinguosdu@gmail.com>

* on_policy_distillation: port from SGLang to vLLM /v1/completions

Follow-up on the test rename in this PR:
test_qwen2.5_0.5B_opd_sglang.py → test_qwen2.5_0.5B_opd_vllm.py.

The test only spawns a vLLM teacher and exercises the OPD pipeline; the
real broken piece was slime/rollout/on_policy_distillation.py, which
PR #18 left in SGLang request/response shape:

  request fields:
    "max_new_tokens": 0          (vLLM: "max_tokens")
    "return_logprob": True       (sglang-only)
    "logprob_start_len": 0       (sglang-only)
  response parsing:
    reward["meta_info"]["input_token_logprobs"]   (sglang shape)

vLLM 0.21 supports the same workflow natively via `prompt_logprobs`:

  request to POST /v1/completions:
    {
        "model": <teacher>,
        "prompt_token_ids": sample.tokens,
        "max_tokens": 1,
        "temperature": 0,
        "prompt_logprobs": 1,
        "logprobs": 0,
        "skip_special_tokens": False,
    }
  response:
    response["choices"][0]["prompt_logprobs"]   # list[dict[int, Logprob] | None]
      where Logprob is {"logprob": float, "rank": int, "decoded_token": str}

References checked against vllm source:
  - reference/vllm/vllm/entrypoints/openai/completion/protocol.py:91
    (request: prompt_logprobs: int | None)
  - reference/vllm/vllm/entrypoints/openai/completion/protocol.py:487
    (response: prompt_logprobs: list[dict[int, Logprob] | None] | None)
  - reference/vllm/vllm/logprobs.py:13
    (Logprob dataclass: logprob/rank/decoded_token)

Implementation notes:

1. JSON serializes int dict keys as strings, so `_logprob_for_token`
   tries both `pos_entry.get(token_id)` and `pos_entry.get(str(token_id))`.

2. `pos_entry` is `None` at position 0 (no prior context) — handled
   explicitly. We also gracefully degrade if a token at position `i` is
   not in the top-1 logprob dict (falls back to 0.0, same as the prior
   sglang code would do).

3. The Logprob dataclass `decoded_token` field is unused; we only read
   `.logprob`. Both dict and `Logprob` shapes are accepted in case the
   server uses a flatter serialization toggle.

4. `args.opd_teacher_model` is the new model-name arg; falls back to
   `args.hf_checkpoint` if not set, mirroring how vime's other rollout
   paths derive the model name.

Smoke-tested `_logprob_for_token` locally:
  - None entry → 0.0
  - int key + dict value → logprob
  - str key (JSON shape) → logprob
  - missing token → 0.0
  - flattened float value → float

Also drops 3 lines from
tests/unit/backends/vllm_utils/test_arguments.py: the
`--sglang-router-ip`/`--sglang-router-port`/`sglang_router_ip` anti-
regression assertions. Once the slim PR #18 lands and sglang is gone
from the runtime, those assertions are vacuous; treating sglang as
non-existent per the project policy.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Canlin Guo <canlinguosdu@gmail.com>

* tests: drop duplicate smoke test updates from PR40

* test(update_weight_from_tensor): drop stale _apply_monkey_patch_torch_reductions patch

The inner ``with patch(f"{MODULE_PATH}._apply_monkey_patch_torch_reductions"):``
context in _run_update suppressed a helper call that PR #48 has since deleted
from update_weight_from_tensor.py (commit 39bf899 on aoshen/align-ipc-rpc-with-slime).
After that PR lands the patched attribute won't exist and this line raises
AttributeError. Remove it now so the test survives PR #48 merge.

The ``sglang_mod.monkey_patch_torch_reductions = MagicMock()`` stub on the
fake sglang module is intentionally kept: on this branch the production code
still imports it via ``from ..sglang import monkey_patch_torch_reductions``
(both update_weight_from_tensor._apply_monkey_patch_torch_reductions on
PR #40's view of main, and hf_weight_iterator_direct.py at module level).
Removing the stub here would break the test on PR #40 alone; it can be
dropped in a follow-up once PR #48 finishes removing every import site.

Tests: ``tests/unit/backends/megatron_utils/update_weight/test_update_weight_from_tensor.py``
all 6 pass with this change applied to gcl/pr18-tests-ci HEAD.

* tests: drop duplicate top-level test_update_weight_from_tensor.py

The 786-line tests/test_update_weight_from_tensor.py is a stale rebase
leftover from the original PR #18 branch — it predates the IPC test
file PR #22 landed at the canonical unit-test path
(tests/unit/backends/megatron_utils/update_weight/test_update_weight_from_tensor.py)
and predates PR #48's single-RPC weight-version contract.

Comparing the two:
* Both stub sys.modules / torch.distributed at module import time, so
  having two files compounds the test-isolation issue Gemini raised
  (PR #40 comment #1).
* Coverage overlaps materially (e.g. test_ipc_init_called_on_first_update_only
  ≈ test_ipc_init_runs_once — same invariant, different wording).
* The nested file is up-to-date with PR #48's RPC contract
  (update_weights_from_tensor.remote(**fields, weight_version=...));
  the top-level file still uses the pre-#48 lifecycle shape and does
  not exercise the coordinator slot fields.
* The nested path matches repo convention: tests/unit/ for mock-only
  unit tests, tests/ top level for e2e scripts.

Closes Gemini comment #1 on PR #40. Gemini comment #2 (the same stub
pattern in the surviving nested file) is a pre-existing issue from
PR #22 / #48 and out of scope for this rename PR — to be addressed
in a follow-up that converts _install_stubs() to an autouse
module-scoped fixture with save/restore.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(vllm_config): use real get_model_url default endpoint /inference/v1/generate

get_model_url defaults to /inference/v1/generate (PR #18), not /v1/completions.
Aligns this test with PR #18's test_vllm_config.py so the two PRs no longer
conflict on this file and the assertion matches the actual runtime default.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Drop on_policy_distillation.py from tests+CI PR (now owned by runtime PR #18)

The OPD vLLM /v1/completions migration is a runtime change; it was folded into
the core-runtime PR (#18). Restore this file to main here so the two PRs no longer
overlap on it. #18 merges first, so this lands via #18.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Drop unit-test files now owned by runtime PR #18

test_vllm_config.py + the plugin_contracts tests are coupled to #18's runtime
rename (they import vllm_config / vllm_rollout, which #18 creates). They live in
#18; remove them here so the two PRs don't overlap. #18 merges first.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: restore vLLM rollout args dropped during sglang→vllm rename

The mechanical sglang→vllm rename dropped several rollout knobs instead of
mapping them to their vLLM equivalents, weakening CI coverage (cuda-graph
capture caps, speculative decoding, expert parallel). Restore them using the
mapping established by the converted production scripts on main
(run-glm4.7-30B-A3B.sh / run-glm5-744B-A40B.sh), verified against vLLM
AsyncEngineArgs:

  --sglang-cuda-graph-max-bs N            -> --vllm-max-cudagraph-capture-size N
  --sglang-cuda-graph-bs a b c            -> --vllm-cudagraph-capture-sizes a b c
  --sglang-ep-size N                      -> --vllm-enable-expert-parallel
  --sglang-speculative-* (eagle)          -> --vllm-speculative-config '{"method":"eagle","num_speculative_tokens":K}'

Also:
- glm4.7 pd: fix --vllm-max-num-seqs (was 8, taken from cuda-graph-max-bs;
  --sglang-max-running-requests was 16) and split out cuda-graph capture.
- fix sglang→rollout mis-renames in temp-file prefixes (→ vllm_*).
- test_vllm_config: rename test_update_weights_default_true →
  test_update_weights_defaults_to_none (it asserts `is None`).

Dropped sglang flags with no vLLM equivalent (enable-dp-lm-head,
moe-dense-tp-size, watchdog-timeout, mamba-scheduler-strategy,
disaggregation-transfer-backend, enable-metrics) stay dropped; PD KV-transfer
is driven by --prefill-num-servers + the --vllm-config prefill/decode topology.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(plugin_contracts): migrate from sglang_rollout to vllm_rollout

The three plugin-contract tests still imported slime.rollout.sglang_rollout
and called install_stubs(with_sglang_router=True), but _shared.install_stubs
already dropped that parameter — so all three failed at collection
(TypeError: unexpected keyword 'with_sglang_router'). Complete the migration:

- install_stubs(with_sglang_router=True, ...) -> install_stubs(...)
- import generate_and_rm / generate_rollout from slime.rollout.vllm_rollout
- default rollout/eval path string -> slime.rollout.vllm_rollout.generate_rollout
  (matches runtime default at slime/utils/arguments.py:233)
- FakeGenerateState: sglang_enable_deterministic_inference ->
  vllm_enable_deterministic_inference, with group_sampling_seeds defaulting to
  None and gated on the flag (mirrors the already-migrated
  tests/unit/rollout/test_vllm_rollout.py).

All 34 plugin-contract cases pass (were 3 collection errors before).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(update_weight_from_tensor): drop stale slime…megatron_utils.sglang mock

The test pre-registered a sys.modules mock for
slime.backends.megatron_utils.sglang (monkey_patch_torch_reductions), left over
from when update_weight_from_tensor imported it. The module under test no longer
imports that module (its real deps are get_gloo_group / HfWeightIteratorBase /
update_weight_from_distributed), so the mock is dead. Removing it makes tests/
and .github/ fully sglang-free. Test still passes (7/7).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [Clean] Remove SGLang runtime code

Rebuilt against current main so the PR contains only the SGLang runtime
removal -- the docs / tests-ci / examples / scripts / docker portions were
split into separate PRs that have since merged.

- Delete dead SGLang server/runtime code: sglang_utils/{arguments,sglang_engine}.py,
  rollout/sglang_rollout.py, the megatron_utils/sglang.py re-export shim, and all
  docker/**/sglang.patch files.
- Rename the rollout config module sglang_utils/sglang_config.py ->
  vllm_utils/vllm_config.py (SglangConfig -> VllmConfig, _resolve_sglang_config ->
  _resolve_vllm_config, --sglang-config -> --vllm-config); inline the
  GPU_MEMORY_TYPE_* constants in rollout.py.
- Add megatron_utils/fp8_helpers.py for the UE8M0 fp8 helpers formerly re-exported
  through the sglang shim; repoint quantizer_fp8 to it.
- Swap sglang_router -> vllm_router in http_utils/wandb_utils; drop the dead
  sglang-router dependency from requirements.txt.
- Finish the SGLang->vLLM rename in the runtime so it is internally consistent and
  matches the tests landing in the tests/CI PR:
  * router args --router-* -> --vllm-router-* (vllm_router_ip/port/timeout);
  * get_model_url reads vllm_model_routers (aligning with rollout.py);
  * --opd-type sglang -> vllm; engine_overrides rename;
  * sglang_enable_deterministic_inference -> vllm_enable_deterministic_inference,
    wired to a real --vllm-enable-deterministic-inference flag (exports
    VLLM_BATCH_INVARIANT=1);
  * consistent_hash session-id routing uses vllm-router's x-session-id header;
  * drop dead trace helper build_sglang_meta_trace_attrs; de-SGLang comments/docstrings.
- Rename test_sglang_config.py -> test_vllm_config.py and de-SGLang the
  plugin-contract tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Address review: finish de-SGLang + fold OPD/router-policy into runtime

- naming: replace residual generic "rollout engine"/"engine"/"comm" wording
  with concrete vLLM (engine_overrides -> vllm_overrides; arguments help text;
  http_utils comments; rollout.py "inference workers"). sglang->vllm is correct,
  sglang->generic is not.
- megatron_to_hf: drop the q_a_proj/kv_a_proj_with_mqa pairing + _cached_tensors
  global. That was sglang-only: sglang's loader torch.cat's both shards within a
  single load_weights call (needs them co-bucketed), whereas vLLM loads each shard
  independently via stacked_params_mapping into fused_qkv_a_proj. Also fix the
  misleading "merge into single fused name" comment.
- docker/Dockerfile: remove now-dead sglang/sglang-router --no-deps stubs + the
  build-time `import sglang` smoke check (slime no longer imports sglang_router).
- OPD: migrate on_policy_distillation.py teacher logprobs to vLLM /v1/completions
  (prompt_logprobs) instead of sglang return_logprob / meta_info.input_token_logprobs.
- routing replay: register --vllm-router-policy (dest=router_policy) so the
  consistent_hash x-session-id session-affinity path is actually wired (was dead).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Review follow-ups: mirror slime vllm_config parsing + restore vLLM process cleanup

- vllm_config.from_yaml: drop the needless `models_raw` intermediate and iterate
  `data["vllm"]` directly, restoring the "Accept both server_groups / legacy
  engine_groups" comment -- mirrors slime's sglang_config.from_yaml line-for-line.
- command_utils.execute_train: re-add a process kill for leftover rollout engines
  as `pkill -9 -f "vllm serve"` (the old `pkill -9 sglang` was dropped with no
  vLLM equivalent), so stale engines don't hold GPUs/ports across runs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Drop test changes from runtime PR; tests live in the tests+CI PR (#40)

The plugin_contracts tests and the test_sglang_config -> test_vllm_config rename
are coupled to the test/CI rename effort and are owned by #40. Restore them to
main here so #18 is purely the SGLang runtime removal. #18 merges first; #40
rebases and re-lands the vLLM test versions.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fp8_helpers: copy SGLang verbatim (fix import crash); pin vLLM deep_gemm env

- fp8_helpers.py: replace the bespoke rewrite with SGLang's exact implementations
  of quant_weight_ue8m0 / transform_scale_ue8m0 and their DeepGEMM helpers
  (per_block_cast_to_fp8, ceil_to_ue8m0, ceil_div, ceil_align, the torch-impl
  packer). deep_gemm is imported lazily inside the functions (as SGLang does), so
  module import no longer requires deep_gemm. This fixes the module-level
  `NameError: _get_tma_aligned_size` that crashed `import megatron_to_hf` on any
  deep_gemm image, and drops the invented sf-stride fixup block that was not in
  upstream. Only should_deepgemm_weight_requant_ue8m0 stays vLLM-adapted
  (is_deep_gemm_e8m0_used) since SGLang's reads SGLang-internal deep_gemm_wrapper.
- vllm_engine.launch_server_process: set VLLM_USE_DEEP_GEMM=1 +
  VLLM_DEEP_GEMM_WARMUP=relax explicitly (setdefault) alongside VLLM_BATCH_INVARIANT,
  replacing SGLang's removed deep_gemm precompile/warmup envs. All vLLM engine env
  now lives in the subprocess env builder (single source of truth).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fp8_helpers: revert to vLLM impl + fix the import NameError; add vLLM trace attrs

- fp8_helpers.py: keep the vLLM-based implementation (uses vllm.utils.deep_gemm,
  consistent with the vLLM runtime) rather than the SGLang verbatim copy. Fix the
  module-level crash: the `try` block referenced `_get_tma_aligned_size` before it
  was bound (the "pre-imported with fallback" import was never written), which
  raised NameError whenever deep_gemm imported successfully -- and NameError is not
  caught by `except ImportError`, so `import megatron_to_hf` crashed on any
  deep_gemm image. Replace the bogus self-assignment with the real import:
  `from vllm.utils.deep_gemm import get_tma_aligned_size as _get_tma_aligned_size`.
- trace_utils/vllm_rollout: add build_vllm_meta_trace_attrs and attach finish_reason
  + token usage to the vllm_inference_generate span (mirrors SGLang's
  build_sglang_meta_trace_attrs; vLLM responses lack the pd_* timing, which lives
  in vLLM's own OTLP traces).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(opd): score teacher via /inference/v1/generate with prompt_logprobs

Move the vllm OPD teacher path off the OpenAI /v1/completions endpoint onto
vime's native /inference/v1/generate (the same endpoint the rollout engines
use), and fix three latent issues:

1. model field: /inference/v1/generate takes `model` as OPTIONAL. Stop
   defaulting to args.hf_checkpoint (the *student* name, which mis-names a
   teacher!=student server). Add --opd-teacher-model; send `model` only when
   set, otherwise omit it (single-model teacher servers use their loaded model).
2. multimodal: the old code sent image_data to a token-only endpoint, which is
   invalid. Raise NotImplementedError until the
   /v1/chat/completions/render -> /inference/v1/generate flow is wired (mirrors
   slime.rollout.vllm_rollout.generate).
3. logprob robustness: read top-level GenerateResponse.prompt_logprobs, assert
   it is present and length-aligned with token_ids, assert the per-sample tensor
   covers response_length, and raise (not silently return 0.0) on a missing
   token logprob. vLLM always includes the actual prompt token in
   prompt_logprobs, so a miss is a real error.

Alignment is unchanged (plp[i] <-> tokens[i], skip pos 0, take [-response_length:]).

Follow-up (separate, in the tests PR): the OPD e2e test must launch a teacher
that exposes /inference/v1/generate and point --rm-url at it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(clean-sglang): purge SGLang from tools, train scripts, and build infra

tools/: drop dead `args.sglang_enable_ep_moe` shim (read nowhere); reword
profile/replay helpers to vLLM and map analyzer hints to vLLM flags
(--enforce-eager, --gpu-memory-utilization). train{,_async}.py: comments
SGLang -> vLLM.

build infra: remove build_conda.sh (SGLang-only conda path); drop the GB300
sgl-kernel install from the Dockerfile; delete docker/npu_patch/ wholesale.

docker base image: bump to vLLM v0.22.0. justfile ARM recipes now pin the real
multi-arch vLLM base images instead of the dead SGLANG_IMAGE_TAG/
ENABLE_SGLANG_PATCH build-args -- cu129-arm64 -> v0.22.0-cu129-ubuntu2404
(CUDA 12.9), cu13-arm64 -> v0.22.0-ubuntu2404 (the default-CUDA tag is already
CUDA 13.0) + ENABLE_CUDA_13=1. vLLM tags are multi-arch manifests, so docker
selects the arm64 image automatically on an ARM host.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* ci(clean-sglang): drop conda-build workflow (ran deleted build_conda.sh on SGLang image)

The single build-conda job ran `bash build_conda.sh` (removed in the previous
commit) inside an lmsysorg/sglang container. With the SGLang-only conda path
gone, the whole workflow is dead.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(clean-sglang): fix stale SGLang refs in docs/skills; tidy comments

docs/conf.py: point the "edit on GitHub" links at vllm-project/vime instead of
the inherited sgl-project.github.io repo. .claude/skills/*: update the dead
`slime/rollout/sglang_rollout.py` references to `vllm_rollout.py` (the real
default is slime.rollout.vllm_rollout.generate_rollout).

justfile: drop the redundant BASE_IMAGE override on release-cu129-arm64 (it
equalled the Dockerfile default; the multi-arch manifest already resolves
arm64). train{,_async}.py: drop stray "the" in the W&B comment.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(tests): use method=mtp (not eagle) in vllm speculative config

The migrated speculative configs pass no draft `model`, so method=eagle
raises "num_speculative_tokens was provided but without speculative model"
in vLLM's SpeculativeConfig. These models carry embedded MTP layers, so
method=mtp is correct and unblocks the mimo MTP-only-grad test (#19).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(cleanup): target renamed vLLM subprocesses in pkill so VRAM is freed

vLLM's set_process_title() renames the VRAM-holding subprocesses
(VLLM::EngineCore, VLLM::Worker_TP*, vllm::router), so their cmdline no
longer contains "vllm serve". The previous `pkill -9 -f "vllm serve"`
matched only the launcher and left engine/worker children holding GPU
memory, leaking it into the next run — masked only by the indiscriminate
`pkill -9 python`, which is unsafe on colocate/shared nodes.

Match both the launcher and the renamed children with
`pkill -9 -f '[v]llm serve|VLL[M]::'`; the [v]/[M] bracket trick keeps the
pattern from matching pkill's own cmdline. This makes the broad python
kill unnecessary, so its already-commented-out lines are removed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(speculative): use method=mtp (not eagle) for embedded-MTP models

vLLM's SpeculativeConfig requires an explicit draft `model` for
method=eagle; with only num_speculative_tokens set it raises
"num_speculative_tokens was provided but without speculative model".
The migrated configs in scripts/examples/docs pass no model, so they must
use method=mtp, which reuses the target checkpoint's embedded MTP layer
(DeepSeek-R1, GLM-4.x-MoE, MiMo, Qwen3-Next/3.5).

The two docs examples that pass an explicit "model" are genuine eagle
usage and are left unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(vllm): launch each rollout engine with its ServerGroup's per-group TP

launch_server_process / _init_normal derived tensor-parallel size and
CUDA_VISIBLE_DEVICES from the global --rollout-num-gpus-per-engine, ignoring
the per-engine num_gpus_per_engine already carried on the VLLMEngine actor.

A ServerGroup configured with num_gpus_per_engine greater than the global
flag (e.g. tp=2) therefore launched as tp=1, while the NCCL weight-sync
rendezvous sized world_size from engine_gpu_counts (the per-group value).
The two disagreed: the trainer waited for a rank the under-sized engine
never started, so init_weight_transfer_engine hung for 300s
("3/4 clients joined") and the job failed.

Honor the per-engine num_gpus_per_engine at launch, falling back to the
global flag when unset (matches the SGLang path and PR #66's
_compute_server_args).

Verified on H200: tests/test_qwen2.5_0.5B_vllm_config_distributed now
launches engine0 tp=2 / engine1 tp=1, update_weights completes in 1.1s
(was a 301s timeout), and rollout+eval proceed.

AI assistance (Claude Code) was used for this change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(ckpt): add --dist-ckpt-optim-fully-reshardable for PAO+offload save/load

test_qwen3_4B_ckpt.py uses precision-aware optimizer + cpu-offload
(HybridDeviceOptimizer). Under the default dp_reshardable (bucket-centric)
optimizer sharding, save/load produce unequal-length param_state lists, so
dist-ckpt load fails with
"Cannot merge two lists with different lengths (81 and 79)".

fully_reshardable is model-centric and immune to bucket-layout changes.
Verified on the r3 image (Megatron-LM 0.16.0rc0 @ 1dcf0da): save+load both
succeed, and source review confirms master_param / step / HybridDeviceOptimizer
sync are handled on this path. This is the flag described in PR #50 that was
never actually merged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(router-args): hybrid naming — vllm_ for ip/port, bare router_ for timeout

vllm-router's RouterArgs.from_cli_args only supports prefix "" or "router_"
(never "vllm_router_"), and excludes host/port from its CLI via
exclude_host_port=True. So:

- --vllm-router-ip / --vllm-router-port keep the vllm_ prefix: RouterArgs does
  not own these CLI flags, vime does (populated via _start_router's manual
  router_args.host/port assignment), so the vllm_ prefix is free and marks them
  as vime-owned endpoint config.
- --router-request-timeout-secs goes bare (dest router_request_timeout_secs): it
  is a genuine RouterArgs field, so it shares the --router-* namespace with
  policy / cache_threshold / retries / … and flows through from_cli_args like
  the other knobs.
- --vllm-router-policy keeps dest=router_policy (unchanged).

Also fixes conftest fixture to seed vllm_router_ip/port (was bare router_ip/port,
which never matched the vllm_engine reader) and updates README/README_zh prose.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Canlin Guo <canlinguosdu@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants