Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
22 commits
Select commit Hold shift + click to select a range
30ee415
sync: update from Slime through #2393
aoshen02 Sep 19, 2026
b7a0762
fix: complete vLLM sync translations
aoshen02 Sep 19, 2026
3f5f82d
sync: complete Slime update through #2394
aoshen02 Sep 21, 2026
cd3992f
Sync Slime through score centering
aoshen02 Sep 24, 2026
cfcf9c8
fix: rebase vLLM patches onto current nightly
aoshen02 Sep 24, 2026
f58399c
fix: remove stale async CI entry and obsolete metrics fallbacks
aoshen02 Sep 28, 2026
72bdd14
feat: transport sampling-mask logprobs for score centering
aoshen02 Sep 28, 2026
30ce1fe
fix(vllm): enable scale-out endpoint for rollout TITO
aoshen02 Sep 28, 2026
3094284
fix(rollout): keep requests during sleep and enable external PD TITO
aoshen02 Sep 28, 2026
c1e10a5
fix(ci): enable TITO on external OPD teacher
aoshen02 Sep 28, 2026
927a686
test: enable scale-out render endpoint in Geo3K e2e
aoshen02 Sep 29, 2026
02825ef
feat: mirror Slime persistent rollout queue and distributed fully async
aoshen02 Sep 29, 2026
651b30b
test: exercise straw transport in mirrored GPU suites
aoshen02 Sep 29, 2026
2733663
style: match repository formatter for straw sync
aoshen02 Sep 29, 2026
a9d013e
ci: install straw in CPU contract container
aoshen02 Sep 29, 2026
e95a5b5
sync: mirror Slime #2427 straw checkpoint rollback
aoshen02 Sep 30, 2026
389da74
ci: install straw dependency for sync tests
aoshen02 Sep 30, 2026
17c155f
test: pin vllm generation config for score centering
aoshen02 Sep 30, 2026
29e0517
fix: bind straw coordinator in vllm rollout
aoshen02 Sep 30, 2026
1ae8210
docs: record full GPU gate procedure
aoshen02 Sep 30, 2026
eaa075d
sync: mirror Slime #2432 Straw queue and replay changes
aoshen02 Sep 30, 2026
3f48e09
ci: quote Straw requirement in Buildkite shell steps
aoshen02 Sep 30, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 9 additions & 5 deletions .buildkite/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,9 +17,9 @@ The four test steps depend on the pre-commit gate. Each suite runs its files
sequentially inside one step because these queues boot a fresh EC2 instance
per job — a per-file matrix would be mostly boot + pip-install time.
Most always-on CPU steps use the standard `python:3.11` image and install their
lightweight dependencies at runtime. `upstream-sync-cpu` uses
`vllm/vime:latest` because the synchronized GLM and checkpoint tests import the
image-pinned Megatron stack even though they do not allocate a GPU.
lightweight dependencies at runtime. `upstream-sync-cpu` uses `VIME_CI_IMAGE`
(defaulting to `vllm/vime:latest`) because the synchronized GLM and checkpoint
tests import the image-pinned Megatron stack even though they do not allocate a GPU.

## Creating the pipeline (one-time, Buildkite UI)

Expand Down Expand Up @@ -61,6 +61,10 @@ the per-test `VIME_TEST_USE_DEEPEP` / `VIME_TEST_USE_FP8_ROLLOUT` /
The block uses `blocked_state: passed`, so a build whose CPU steps are green
reports a passing commit status even if nobody unblocks the GPU gate.

For sync validation, always select all six suites. The REST unblock payload
passes the selection directly, for example `{"gpu-suites": "short\nvllm-config\nmegatron\nvime-customized\nprecision\nckpt"}`;
do not wrap it in a second `fields` object.

GPU jobs run on the shared **`mithril-h100-pool`** queue, following the same
pattern vllm-omni uses for it: each job is a Kubernetes pod (agent-stack-k8s
`kubernetes` plugin) on an H100 SXM node, with GPUs allocated via
Expand All @@ -70,8 +74,8 @@ startup, so a warm HF cache is all they need. `WANDB_API_KEY` is not wired up
yet; runs report without wandb until it's added (e.g. as a k8s secret in the
pod spec).

GPU jobs use `vllm/vime:latest`. Rebuild and publish that image before validating
Dockerfile or vLLM patch changes.
Set `VIME_CI_IMAGE` to an immutable candidate digest for image-backed jobs;
otherwise they use `vllm/vime:latest`. Do not update `latest` before merge.

## Keeping it in sync

Expand Down
5 changes: 3 additions & 2 deletions .buildkite/gpu_suites.py
Original file line number Diff line number Diff line change
Expand Up @@ -23,15 +23,14 @@
import subprocess

GPU_QUEUE = "mithril-h100-pool"
CI_IMAGE = "vllm/vime:latest"
CI_IMAGE = os.environ.get("VIME_CI_IMAGE", "vllm/vime:latest")
HF_CACHE_HOST_PATH = "/mnt/hf-cache"
HF_HOME = "/root/.cache/huggingface"
NODE_INSTANCE_TYPE = "gpu-h100-sxm"

# (test_file, num_gpus, extra_args, env overrides)
SUITES = {
"short": [
("test_qwen3.5_0.8B_gsm8k_async_short.py", 4, "", {}),
("test_qwen3.5_0.8B_gsm8k_short.py", 4, "", {}),
("test_qwen2.5_0.5B_fully_async_short.py", 4, "", {}),
],
Expand Down Expand Up @@ -64,8 +63,10 @@
("test_moonlight_16B_A3B_r3.py", 8, "", {"ENABLE_EVAL": "0"}),
("test_mimo_7B_mtp_only_grad.py", 8, "", {}),
("test_qwen2.5_0.5B_debug_rollout_then_train.py", 8, "", {}),
("test_straw_checkpoint_fork.py", 4, "", {}),
("test_qwen2.5_0.5B_opd_vllm.py", 8, "", {}),
("test_qwen2.5_0.5B_fanout_short.py", 4, "", {}),
("test_qwen2.5_0.5B_score_centering.py", 2, "", {}),
("test_qwen2.5_0.5B_debug_train_dump_e2e.py", 8, "", {}),
("test_qwen3_4B_external_pd.py", 6, "", {"VIME_TEST_UPDATE_MODE": "delta"}),
],
Expand Down
22 changes: 19 additions & 3 deletions .buildkite/pipeline.yml
Original file line number Diff line number Diff line change
Expand Up @@ -64,9 +64,10 @@ steps:
python:3.11 bash -c '
set -euo pipefail
pip install -q torch --index-url https://download.pytorch.org/whl/cpu
pip install -q pytest numpy packaging pyyaml omegaconf tqdm httpx requests ray pybase64 pylatexenc sympy aiohttp pillow safetensors transformers cloudpickle blake3 xxhash zstandard psutil wandb
pip install -q pytest numpy packaging pyyaml omegaconf tqdm httpx requests ray pybase64 pylatexenc sympy aiohttp pillow safetensors transformers cloudpickle blake3 xxhash zstandard psutil wandb "straw-queue>=0.1.2"
pip install -q -e . --no-deps
python tests/test_megatron_argument_validation.py
python tests/test_optional_straw.py
python tests/test_value_temperature.py
python tests/test_docs_consistency.py
python tests/test_placement_group.py
Expand Down Expand Up @@ -140,18 +141,26 @@ steps:
-e GIT_CONFIG_PARAMETERS="'safe.directory=/workspace'" \
-e GLOO_SOCKET_IFNAME=lo -e TP_SOCKET_IFNAME=lo \
-v "$$PWD:/workspace" -w /workspace \
vllm/vime:latest bash -lc '
"$${VIME_CI_IMAGE:-vllm/vime:latest}" bash -lc '
set -euo pipefail
pip install -q "straw-queue>=0.1.2" --break-system-packages
pip install -q -e . --no-deps --break-system-packages
for test_file in \
tests/test_advantage_whiten_cp.py \
tests/test_accelerator.py \
tests/test_block_fp8_zero_block.py \
tests/test_cuda_stack_offload.py \
tests/test_deep_ep_tms_patch.py \
tests/test_distributed_rollout.py \
tests/test_discounted_returns.py \
tests/test_eval_config.py \
tests/test_filter_long_prompt.py \
tests/test_fully_async_rollout.py \
tests/test_queue_sample_codec.py \
tests/test_rollout_straw_transport.py \
tests/test_routed_experts_disk_layout.py \
tests/test_straw_fully_async_recovery.py \
tests/test_straw_r3.py \
tests/test_glm52_layerwise_comparison.py \
tests/test_glm5_indexer_q_norm.py \
tests/test_glm5_indexer_short_context.py \
Expand All @@ -160,11 +169,17 @@ steps:
tests/test_model_provider_freeze.py \
tests/test_policy_loss.py \
tests/test_ppo_kl_metric.py \
tests/test_published_payload_retry.py \
tests/test_rollout_metadata_index.py \
tests/test_score_centering.py \
tests/test_score_centering_transport.py \
tests/test_process_rollout_data.py \
tests/test_qwen3_5_vl_native.py \
tests/test_qwen3_linear_attention_cu_seqlens.py \
tests/test_read_file_slicing.py \
tests/test_reloadable_process_group_world.py \
tests/test_r3_spill.py \
tests/test_rollout_buffer_order.py \
tests/test_rollout_data_utils.py \
tests/test_rollout_metrics.py \
tests/test_rollout_sample_hooks.py \
Expand All @@ -174,6 +189,7 @@ steps:
tests/test_tau_bench_token_delta.py \
tests/test_train_data_utils.py \
tests/test_update_weight_factory.py \
tests/test_vllm_server_control.py \
tests/observability/test_trace_utils.py; do
python -m pytest "$$test_file"
done
Expand Down Expand Up @@ -222,7 +238,7 @@ steps:
value: short
- label: "run-ci-vllm-config — 8 GPU, 4 tests"
value: vllm-config
- label: "run-ci-megatron — up to 8 GPU, 21 runs"
- label: "run-ci-megatron — up to 8 GPU, 23 runs"
value: megatron
- label: "run-ci-vime-customized — 1–8 GPU, 5 tests"
value: vime-customized
Expand Down
4 changes: 2 additions & 2 deletions .claude/skills/add-dynamic-filter/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,7 @@ from vime.rollout.filter_hub.base_types import DynamicFilterOutput
return DynamicFilterOutput(keep=True, reason=None)
```

Buffer filter (called in `vime/rollout/data_source.py`):
Buffer filter (called in `vime/data/data_source.py`):

```python
def buffer_filter(args, rollout_id, buffer, num_samples):
Expand Down Expand Up @@ -98,4 +98,4 @@ Example wiring:
- Dynamic filter types: `vime/rollout/filter_hub/base_types.py`
- Dynamic filter example: `vime/rollout/filter_hub/dynamic_sampling_filters.py`
- Rollout generation hook points: `vime/rollout/vllm_rollout.py`
- Buffer filter hook point: `vime/rollout/data_source.py`
- Buffer filter hook point: `vime/data/data_source.py`
7 changes: 7 additions & 0 deletions .claude/skills/vime-code-review-preferences/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,3 +27,10 @@ Apply these lightweight review heuristics when changing vime code.
- Add an abstraction only when it removes real duplication, hides fragile mechanics, or clarifies ownership.
- When a review comment points out repeated indirection, look for a smaller public surface rather than adding another alias.
- Preserve existing behavior intentionally. If cleanup changes error handling, logging, or failure visibility, call that out in the final response.

## Keep Runtime State Out of Args

- Treat `args` as configuration. Do not add ad hoc `args._xxx` attributes to pass actor handles, caches, recovery progress, current versions, or other runtime state between components. Renaming these to public attributes does not fix the ownership problem.
- Put live state on the component that owns its lifecycle. Pass dependencies explicitly, and use typed return values for startup plans or restoration results shared between components.
- Avoid replacing scattered temporary fields with a generic `args.context` bag. Keep each state object's responsibility and lifetime clear.
- Preserve existing custom hook signatures, call arguments, and return conventions when refactoring internal state. User implementations must not need new parameters merely to accommodate an internal cleanup.
2 changes: 2 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,6 +53,8 @@ The vLLM community horizontally supports many LLM post-training frameworks, incl
- **rollout (vLLM + router)**: Launches vLLM inference engines and routes generation requests; custom generate functions can wrap generation with multi-turn loops, tool calls, environment/sandbox interaction, and verifier-based rewards.
- **data buffer**: A bridge module that manages prompt initialization, custom data, and rollout generation methods, including agentic workflows that produce samples through the same interface.

The default payload transport is Ray `object-store`. `--rollout-data-transport straw` persists prompt tasks, rollout continuations, and training batches on shared storage. See the [straw architecture and recovery guide](docs/en/advanced/straw.md).

## Quick Start

For a comprehensive quick start guide covering environment setup, data preparation, training startup, and key code analysis, please refer to:
Expand Down
2 changes: 2 additions & 0 deletions README_zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,6 +53,8 @@ vLLM 社区横向支持许多 LLM post-training 框架,包括(按字母顺
- **rollout (vLLM + router)**:启动 vLLM 推理引擎并路由生成请求;自定义生成函数可以在其上封装多轮循环、工具调用、环境/沙盒交互和基于 verifier 的奖励;
- **data buffer**:桥梁模块,管理 prompt 初始化、自定义数据与 rollout 生成方法,包括通过同一接口产出样本的 agent 工作流。

默认载荷传输为 Ray `object-store`。`--rollout-data-transport straw` 可在共享存储上持久化 prompt 任务、rollout continuation 和训练 batch。详见 [straw 架构与恢复指南](docs/zh/advanced/straw.md)。

## 快速开始

有关环境配置、数据准备、训练启动和关键代码分析的完整快速开始指南,请参考:
Expand Down
3 changes: 3 additions & 0 deletions docker/Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -153,15 +153,18 @@ RUN cd Megatron-LM && \
# so apply the independently maintained patches in their validated order.
COPY docker/patch/${PATCH_VERSION}/vllm-pull_weights.patch /tmp/vllm-pull_weights.patch
COPY docker/patch/${PATCH_VERSION}/vllm.patch /tmp/vllm.patch
COPY docker/patch/${PATCH_VERSION}/vllm-score-centering.patch /tmp/vllm-score-centering.patch
COPY docker/patch/${PATCH_VERSION}/vllm-pd-request-metrics.patch /tmp/vllm-pd-request-metrics.patch
COPY docker/patch/${PATCH_VERSION}/vllm-inflight-queue-diagnostics.patch /tmp/vllm-inflight-queue-diagnostics.patch
RUN VLLM_SITE="$(python3 -c 'import os, vllm; print(os.path.dirname(os.path.dirname(vllm.__file__)))')" && \
cd "$VLLM_SITE" && \
git apply -v /tmp/vllm-pull_weights.patch && \
git apply -v --allow-empty /tmp/vllm.patch && \
git apply -v /tmp/vllm-score-centering.patch && \
git apply -v /tmp/vllm-pd-request-metrics.patch && \
git apply -v /tmp/vllm-inflight-queue-diagnostics.patch && \
rm /tmp/vllm-pull_weights.patch /tmp/vllm.patch \
/tmp/vllm-score-centering.patch \
/tmp/vllm-pd-request-metrics.patch \
/tmp/vllm-inflight-queue-diagnostics.patch

Expand Down
33 changes: 33 additions & 0 deletions docker/patch/latest/megatron.patch
Original file line number Diff line number Diff line change
Expand Up @@ -1005,3 +1005,36 @@ index e9736ac08..6567ed426 100644
)
for (model_chunk_idx, model_chunk) in enumerate(model)
]
diff --git a/megatron/training/checkpointing.py b/megatron/training/checkpointing.py
index 7c87eca191a525915502a1deb7622f03990c89de..bdb62df91b9ca2fe6e13d166a21f38e1b1e90e64 100644
--- a/megatron/training/checkpointing.py
+++ b/megatron/training/checkpointing.py
@@ -212,8 +212,8 @@
iteration, release = read_metadata(tracker_filename)

# Allow user to specify the loaded iteration.
- if getattr(args, "ckpt_step", None):
- iteration = args.ckpt_step
+ if getattr(args, "ckpt_step", None) is not None:
+ iteration, release = args.ckpt_step, False

return get_checkpoint_name(load_dir, iteration, release, return_base_dir=True)

@@ -1185,14 +1185,14 @@
iteration, release = read_metadata(tracker_filename)

# Allow user to specify the loaded iteration.
- if getattr(args, "ckpt_step", None):
- iteration = args.ckpt_step
+ if getattr(args, "ckpt_step", None) is not None:
+ iteration, release = args.ckpt_step, False

# Record the iteration loaded (stored separately from args to avoid
# polluting checkpoints, since args is saved in checkpoints).
set_loaded_iteration(iteration)

- if non_persistent_iteration != -1: # there is a non-persistent checkpoint
+ if non_persistent_iteration != -1 and getattr(args, "ckpt_step", None) is None:
if non_persistent_iteration >= iteration:
return _load_non_persistent_base_checkpoint(
non_persistent_global_dir,
Loading