Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
158 commits
Select commit Hold shift + click to select a range
08ce80b
temp save rfc
SamitHuang Mar 2, 2026
3af806e
add plan
SamitHuang Mar 3, 2026
48fbde3
update
SamitHuang Mar 3, 2026
d78bb43
[docker] remove true on policy patches (#1661)
zhuzilin Mar 3, 2026
0104a9e
[fix]: Qwen3.5-35B-A3B 8-GPU: set TP size to 2 for num_query_groups=2…
none0663 Mar 4, 2026
e4faf63
Remove FSDP support (#1664)
zhuzilin Mar 4, 2026
988c45a
docs: add OpenClaw-RL to projects built upon slime (#1635)
yinjjiew Mar 4, 2026
f8ceed6
qwen2.5 0.5b non-colocate (first attempt ok, but nccl error later)
SamitHuang Mar 4, 2026
2caa4a0
add convert script
SamitHuang Mar 4, 2026
8caa8ba
add setup doc
SamitHuang Mar 4, 2026
de84e10
Support setting update weights in sglang_config (#1665)
zhuzilin Mar 4, 2026
25ee005
fix nccl error by NcclBridge subprocess
SamitHuang Mar 5, 2026
09f534a
Add rollout backend client and test qwen2.5-0.5b non-colocate training
SamitHuang Mar 5, 2026
ab7eb0b
eliminate gpu to cpu weight transfer
SamitHuang Mar 5, 2026
411e2d2
Eliminate intermediate CPU tensors for faster weight transfer
SamitHuang Mar 5, 2026
546d2ad
Revise weight synchronization strategy in goal plan
SamitHuang Mar 5, 2026
dd6888d
[fix] Fix numerical accuracy issue in dynamic sampling filter (#1674)
Django-Jiang Mar 7, 2026
2cd28d6
sync from internal (#1677)
zhuzilin Mar 7, 2026
9268231
bugfixes from community (#1678)
zhuzilin Mar 7, 2026
450e9f3
Fix: pass return_tensors in text_kwargs for transformers>=5.0.0 compa…
coding-famer Mar 7, 2026
ce204d1
Fix missing packed_seq_params in bshd qkv_format (#1649)
coding-famer Mar 7, 2026
241c75b
[Multimodal][Model] Qwen3.5 VL training example/support (#1676)
coding-famer Mar 7, 2026
ceea3b0
update docs (#1680)
zhuzilin Mar 8, 2026
b2c5fc7
update docs (#1681)
zhuzilin Mar 8, 2026
1018387
support offloading non-updatable server (#1668)
zhuzilin Mar 8, 2026
29e1487
bugfix (#1685)
zhuzilin Mar 8, 2026
22915a3
fix: handle Qwen3.5 in quantize_params_fp8 (#1683)
lawrence-harmonic Mar 8, 2026
80d4bc3
bugfix (#1687)
zhuzilin Mar 8, 2026
1d31a49
Fix Qwen3.5 & Qwen3-Next linear attention cu_seqlens missing (#1686)
huang3eng Mar 8, 2026
a4492dd
fix: use semantic version comparison for PyTorch >= 2.6 detection (#1…
abatilo Mar 8, 2026
cc9e3bb
[Fix] Minor fix for properly finishing / flushing wandb logging metri…
silunw Mar 8, 2026
750d977
Autofix/issue 1578 hf2megatron arg suffix (#1636)
yitianlian Mar 8, 2026
b749f70
bugfix (#1688)
zhuzilin Mar 8, 2026
77b173b
fix(examples): update strands_sglang example to v0.3.x API (#1684)
Lawhy Mar 8, 2026
2640e6c
[docker] cherry pick qwen3.5 bugfix (#1691)
zhuzilin Mar 9, 2026
b843ac1
bugfix/fix Qwen3.5 dense model precision bug in TP_SIZE>1 from sglang…
mzusman Mar 11, 2026
5182e2d
Fix/qwen3 5 mtp bridge (#1702)
huang3eng Mar 11, 2026
6f11313
support epd for glm4.6v (#1704)
hanwen-sun Mar 11, 2026
8d9378e
[docker] support epd for glm4.6v (#1707)
zhuzilin Mar 11, 2026
dcfc790
remove script
invalid-email-address Mar 11, 2026
57e7668
[docker] store v0.5.9 patch (#1710)
zhuzilin Mar 11, 2026
e70e647
Add GLM-4.7-Flash MTP training support (#1712)
zhuzilin Mar 11, 2026
6195417
[release] bump to v0.2.3 (#1682)
zhuzilin Mar 12, 2026
6006246
feat: add GLM-4.6V MoE VL bridge with CP support (#1715)
zhuzilin Mar 12, 2026
ae2590a
fix: resolve rope_theta from rope_parameters dict in HF config valida…
zhuzilin Mar 13, 2026
1ad13f7
[docker] patches for glm4.6v, kimi k2.5 and dsa cp only (#1722)
zhuzilin Mar 13, 2026
08b201b
[docker] support IndexCache
invalid-email-address Mar 13, 2026
183e525
Fix CUDA IPC cache leaks during weight updates (#1731)
zhuzilin Mar 17, 2026
6e3699c
[docker] update megatron (#1729)
zhuzilin Mar 18, 2026
02fef7e
[docker] Fix IndexCache with mla model (#1736)
zhuzilin Mar 18, 2026
e4d22dc
[slime-router] support pd disaggregation and remove radix tree middle…
zhuzilin Mar 18, 2026
f71f710
Fix glm4v megatron bridge (#1738)
zhuzilin Mar 18, 2026
a1ee76c
[docker] update sglang patch (#1743)
zhuzilin Mar 20, 2026
5996c68
feat: GLM4V multimodal support improvements (#1745)
zhuzilin Mar 21, 2026
a961443
feat: placeholder worker type, metrics router, and GPQA letter range …
zhuzilin Mar 21, 2026
051e91a
always enable_metrics and remove dp context (#1747)
zhuzilin Mar 21, 2026
478b807
fix: resolve SP/CP gradient inflation in FLA (linear attention) layer…
zhuzilin Mar 22, 2026
e8e4b64
Update MTP example configs, rename GLM-4.5 to GLM-4.7, clean scripts …
zhuzilin Mar 22, 2026
7f2a03b
Support qwen3.5 loss mask for multi-turn SFT (#1742)
huang3eng Mar 22, 2026
73a1f4d
fix: propagate moe_token_dispatcher_type in bridge model provider (#1…
nanjiangwill Mar 22, 2026
d269838
fix: resolve rope_theta from rope_parameters in DeepseekV32Bridge (#1…
stevewx Mar 22, 2026
28b5e72
chore: translate remaining Chinese comments to English (#1726)
WangHong-yang Mar 22, 2026
e46e660
feat: add Qwen3.5-4B model support (#1721)
shihaohou Mar 22, 2026
5fa6755
fix: http_utils. disable system proxy for internal SGLang httpx clien…
DongzhuoranZhou Mar 22, 2026
aeaa5b3
fix: auto-detect GPUs in qwen3-4b script (#1700)
ailuntz Mar 22, 2026
10f21d8
fix: quote `$MOE_LAYER_FREQ` (#1689)
lawrence-harmonic Mar 22, 2026
8a83f94
disable router health_check and allow prompt_data is None (#1751)
zhuzilin Mar 23, 2026
d480da0
Router for vllm (#5)
knlnguyen1802 Mar 23, 2026
243d088
small fix on qwen3-235b-a22b launch script (#1719)
Zhuohao-Li Mar 23, 2026
d4c4d3f
sync internal bugfix (#1765)
zhuzilin Mar 25, 2026
b82193f
Fix uploading sglang metrics to wandb (#1768)
zhuzilin Mar 26, 2026
5df168b
use zhuzilin/sgl-router for sglang-router (#1770)
zhuzilin Mar 26, 2026
f3a86d5
[docker] update sgl-router (#1772)
zhuzilin Mar 27, 2026
c291ef8
[Multimodal] Add Multimodal OPD support (#1760)
coding-famer Mar 27, 2026
dbd4e73
refactor: remove slime router (#1773)
zhuzilin Mar 27, 2026
0988f0f
Add rollout trace timeline viewer (#1776)
zhuzilin Mar 28, 2026
f9f7b56
[Fix] Fix duplicate Megatron LR scheduler resume when optimizer state…
kaysonyu Mar 29, 2026
353ac7b
Support FP8 conversion for Qwen3.5 (#1769)
peterjc123 Mar 29, 2026
64e1e68
fix typo (#1759)
albaNnaksqr Mar 29, 2026
6f70479
[Fix]Fix some bugs/clean up (#1756)
coding-famer Mar 29, 2026
f96b71f
(fix):not have encoder_only attr cause run failed (#1741)
wangyufak Mar 29, 2026
ec176ca
update docs
invalid-email-address Mar 29, 2026
f2270ad
remove redundant envvar
invalid-email-address Mar 29, 2026
bd217a6
some minor cleanup
zhuzilin Mar 29, 2026
5d7ca97
[release] bump to v0.2.4 (#1777)
zhuzilin Mar 29, 2026
c154078
Plan refactor vllm/sglang
knlnguyen1802 Mar 30, 2026
1a2dcf5
Code implemented
knlnguyen1802 Mar 31, 2026
8addb37
Fix bug
knlnguyen1802 Mar 31, 2026
9689a97
Fix bug
knlnguyen1802 Mar 31, 2026
3753432
Fix bug
knlnguyen1802 Mar 31, 2026
f1e7554
Fix port
knlnguyen1802 Mar 31, 2026
91cc780
Fix config
knlnguyen1802 Mar 31, 2026
be1ecd4
Fix bug MOE weight sync
knlnguyen1802 Apr 1, 2026
e7216d8
Fix bug vllm transfer weight
knlnguyen1802 Apr 1, 2026
2343647
Fix weight sync
knlnguyen1802 Apr 1, 2026
8498b7b
Fix
knlnguyen1802 Apr 1, 2026
8a41184
Fix config
knlnguyen1802 Apr 2, 2026
acc9690
Change name config
knlnguyen1802 Apr 2, 2026
3e765c8
pass critic role through to create RayTrainGroup (#1797)
znculee Apr 3, 2026
43c5174
fix qwen3.5 397B converting error when enable expert parallel (#1799)
xutianming Apr 3, 2026
887fba1
fix(geo3k-vlm-sft): remove --apply-chat-template from SFT launch scri…
DongzhuoranZhou Apr 3, 2026
b4b8fa9
Add host memory metrics to available_memory function (#1764)
peterjc123 Apr 3, 2026
2a89c16
[WIP] fix loss oom (#1788)
lilei199908 Apr 4, 2026
1eed249
sync from internal (#1805)
zhuzilin Apr 5, 2026
ccf0421
sync from internal (#1807)
zhuzilin Apr 5, 2026
2432b3d
feat: add npu patch for qwen3-vl-8b grpo & ppo (#1750)
Apr 7, 2026
ce76977
fix missing position_ids in log-prob forward step (#1809)
znculee Apr 7, 2026
31c99b2
feat: add support for including missing weights from origin HF checkp…
peterjc123 Apr 7, 2026
8533c23
[Fix] Initialize grad_norm before found_inf skip path (#1762)
kaysonyu Apr 7, 2026
913b3c4
[conda] Add install custom sgl-router to build_conda.sh (#1813)
zhuzilin Apr 7, 2026
2928bba
Revert no_grad for entropy to prevent comm stuck in dsa (#1822)
zhuzilin Apr 9, 2026
286750a
Add fallback for get_seqlen_balanced_partitions (#1823)
zhuzilin Apr 9, 2026
90934e6
Resolve review
knlnguyen1802 Apr 13, 2026
6b4c373
Try colocated vllm weight
knlnguyen1802 Apr 15, 2026
08ed342
docs: add Relax to notable projects in README (#1834)
Yangruipis Apr 15, 2026
d649cfd
Bugfix: use cpu instead of cuda in convert_torch_dist_to_hf.py when -…
coding-famer Apr 15, 2026
8efb116
[fix] eval sample logging when sample is a list (#1836)
mathewjhan Apr 16, 2026
49d760f
[Draft] Local runable dev
knlnguyen1802 Apr 17, 2026
7d6b2e0
[Fix] Fix cuda-python pin in build_conda.sh (#1827)
kaysonyu Apr 20, 2026
5b688aa
fix entropy bug and update code (#1846)
lilei199908 Apr 20, 2026
8acc034
Revert "Add fallback for get_seqlen_balanced_partitions" (#1848)
zhuzilin Apr 21, 2026
79342be
fix (#1849)
lilei199908 Apr 21, 2026
ee2702b
Fix offload train
knlnguyen1802 Apr 21, 2026
f44af13
Add support for NVIDIA DGX Spark (GB10 / sm_121a, arm64) (#1835)
boots-coder Apr 21, 2026
cb77fc5
Fix offload train
knlnguyen1802 Apr 22, 2026
ba06919
Fix offload_rollout
knlnguyen1802 Apr 22, 2026
3491aae
Fix vllm offload
knlnguyen1802 Apr 24, 2026
af2bfb4
Fix offload traing
knlnguyen1802 Apr 24, 2026
14aa09a
Fix offload weight
knlnguyen1802 Apr 24, 2026
d82d1d0
Fix offload weight
knlnguyen1802 Apr 24, 2026
75af529
refactor/ppo (#1856)
lilei199908 Apr 24, 2026
ea28a9a
[docker] cleanup sglang patch (#1859)
zhuzilin Apr 25, 2026
3cdec0c
[docker] update v0.5.9 patch
zhuzilin Apr 25, 2026
f65c6e8
Rename critic config to megatron config (#1866)
zhuzilin Apr 27, 2026
5d41cf7
[Fix] Use Ray ObjectRef await instead of asyncio.to_thread in distrib…
ryang-max Apr 28, 2026
27b899f
chore: include length context in slice_log_prob_with_cp assert (#1862)
leofan-i Apr 28, 2026
f3e7bd7
[docker] upgrade megatron to 1dcf0dafa (#1867)
zhuzilin Apr 28, 2026
9a022d8
fix ppo value head load bugs (#1878)
lilei199908 Apr 29, 2026
f588add
[docker] upgrade sglang to v0.5.10.post1 (#1874)
zhuzilin Apr 30, 2026
07beb18
[docs] update docs
zhuzilin Apr 30, 2026
fe1152b
[docker] update megatron-bridge and add qwen3.6 tests (#1884)
zhuzilin Apr 30, 2026
9b50665
fix lint
zhuzilin Apr 30, 2026
16924b6
Fix(checkpoint): add resume/pause in save_model() for offload_train (…
Procrastinatorrrr May 6, 2026
1027409
fix ppo value offload bugs (#1882)
lilei199908 May 6, 2026
3e0a3ca
fix qwen3.6 hf config validation bug (#1889)
zhuzilin May 6, 2026
4cacab3
Add missing metrics to log (#1890)
zhuzilin May 6, 2026
04059e5
fix(qwen3_next): use torch.get_default_dtype() — get_current_dtype do…
HeatherLiuzh May 6, 2026
477541d
Fix location error in install script (#1877)
selfanti May 6, 2026
82007fa
Only allow --allgather-cp for DSA model (#1891)
zhuzilin May 6, 2026
bf9b1a3
Migrate internal feature (#1897)
zhuzilin May 9, 2026
8ef1fb4
[Fix] Fix distributed POST actor concurrency split (#1880)
kaysonyu May 9, 2026
c8aaf01
Fix CI: update rollout_data_postprocess plugin contract for new call …
jingshenghang May 11, 2026
5b326e6
Patch Megatron TP grad coalesce to chunked all-reduce (#1899)
jingshenghang May 11, 2026
41dc3b6
fix: harden retool rollout against multi-turn / retry desync (#1861)
leofan-i May 11, 2026
a7a3ee1
Fix log file
knlnguyen1802 May 14, 2026
e7dbc8d
Rebase
knlnguyen1802 May 15, 2026
b015d72
Fix import engine group
knlnguyen1802 May 15, 2026
26f979d
Fix rebase code
knlnguyen1802 May 15, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/conda-ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ jobs:
runs-on: self-hosted
container:
image: lmsysorg/sglang:v0.5.0rc0-cu126
options: --gpus all --ipc=host --shm-size=16g --ulimit memlock=-1 --ulimit stack=67108864 --memory=0 --memory-swap=0 -v /mnt/nvme0n1/models:/root/models -v /mnt/nvme0n1/datasets:/root/datasets
options: --privileged --cap-add SYS_NICE --security-opt seccomp=unconfined --gpus all --ipc=host --shm-size=16g --ulimit memlock=-1 --ulimit stack=67108864 --memory=0 --memory-swap=0 -v /mnt/nvme0n1/models:/root/models -v /mnt/nvme0n1/datasets:/root/datasets

defaults:
run:
Expand Down
654 changes: 410 additions & 244 deletions .github/workflows/pr-test.yml

Large diffs are not rendered by default.

226 changes: 149 additions & 77 deletions .github/workflows/pr-test.yml.j2
Original file line number Diff line number Diff line change
Expand Up @@ -2,28 +2,31 @@
'e2e-test-short': {
'label': 'run-ci-short',
'tests': [
{'test_file': 'test_qwen2.5_0.5B_gsm8k_async_short.py', 'num_gpus': 4},
{'test_file': 'test_qwen2.5_0.5B_gsm8k_short.py', 'num_gpus': 4},
{'test_file': 'test_qwen2.5_0.5B_sglang_config.py', 'num_gpus': 8},
{'test_file': 'test_qwen2.5_0.5B_sglang_config_distributed.py', 'num_gpus': 8},
{'test_file': 'test_qwen3.5_0.8B_gsm8k_async_short.py', 'num_gpus': 4},
{'test_file': 'test_qwen3.5_0.8B_gsm8k_short.py', 'num_gpus': 4},
{'test_file': 'test_qwen2.5_0.5B_ppo_critic_only_short.py', 'num_gpus': 4},
],
},
'e2e-test-fsdp': {
'label': 'run-ci-fsdp',
'e2e-test-sglang-config': {
'label': 'run-ci-sglang-config',
'tests': [
{'test_file': 'test_qwen3_4B_fsdp_true_on_policy.py --colocated', 'num_gpus': 4},
{'test_file': 'test_qwen3_vl_4B_fsdp.py', 'num_gpus': 8},
{'test_file': 'test_qwen3_0.6B_megatron_fsdp_align.py', 'num_gpus': 4},
{'test_file': 'test_qwen2.5_0.5B_sglang_config.py', 'num_gpus': 8},
{'test_file': 'test_qwen2.5_0.5B_sglang_config_distributed.py', 'num_gpus': 8},
{'test_file': 'test_sglang_config_mixed_offload.py', 'num_gpus': 8},
{'test_file': 'test_sglang_config_mixed_offload_ft.py', 'num_gpus': 8},
],
},
'e2e-test-megatron': {
'label': 'run-ci-megatron',
'tests': [
{'test_file': 'test_quick_start_glm4_9B.py', 'num_gpus': 8},
{'test_file': 'test_glm4.7_30B_A3B_pd_mooncake.py', 'num_gpus': 8},
{'test_file': 'test_qwen3_30B_A3B.py', 'num_gpus': 8, 'use_deepep': '1', 'use_fp8_rollout': '1'},
{'test_file': 'test_qwen3.6_35B_A3B_pd_mooncake.py', 'num_gpus': 8, 'use_deepep': '1'},
{'test_file': 'test_qwen3_30B_A3B_r3.py', 'num_gpus': 8, 'use_deepep': '1', 'use_fp8_rollout': '1', 'enable_eval': '0'},
{'test_file': 'test_qwen3_30B_A3B_r3.py', 'num_gpus': 8, 'enable_eval': '0'},
{'test_file': 'test_qwen3_4B_ppo.py', 'num_gpus': 8},
{'test_file': 'test_qwen3_4B_ppo_disaggregate.py', 'num_gpus': 8},
{'test_file': 'test_qwen3_4B_ppo_train_critic_only.py', 'num_gpus': 8},
{'test_file': 'test_moonlight_16B_A3B.py', 'num_gpus': 8},
{'test_file': 'test_moonlight_16B_A3B_r3.py', 'num_gpus': 8, 'enable_eval': '0'},
Expand All @@ -36,21 +39,22 @@
'label': 'run-ci-precision',
'tests': [
{'test_file': 'test_qwen3_0.6B_parallel_check.py', 'num_gpus': 8},
{'test_file': 'test_qwen3_0.6B_megatron_fsdp_align.py', 'num_gpus': 4},
],
},
'e2e-test-ckpt': {
'label': 'run-ci-ckpt',
'tests': [
{'test_file': 'test_qwen3_4B_ckpt.py', 'num_gpus': 8},
{'test_file': 'test_qwen3_4B_ckpt.py --async-save', 'num_gpus': 8},
{'test_file': 'test_qwen3_4B_ckpt.py', 'test_args': '--async-save', 'num_gpus': 8},
],
},

'e2e-test-plugin-contracts': {
'label': 'run-ci-plugin-contracts',
'always': True,
'cpu': True,
'tests': [
{'test_file': 'test_megatron_argument_validation.py', 'num_gpus': 0},
{'test_file': 'plugin_contracts/test_plugin_rollout_contracts.py', 'num_gpus': 0},
{'test_file': 'plugin_contracts/test_plugin_runtime_hook_contracts.py', 'num_gpus': 0},
{'test_file': 'plugin_contracts/test_plugin_path_loading_contracts.py', 'num_gpus': 0},
Expand All @@ -62,19 +66,18 @@
'label': 'run-ci-image',
'image': 'slimerl/slime-test:latest',
'tests': [
{'test_file': 'test_qwen2.5_0.5B_gsm8k_async_short.py', 'num_gpus': 4},
{'test_file': 'test_qwen2.5_0.5B_gsm8k_short.py', 'num_gpus': 4},
{'test_file': 'test_qwen3_4B_fsdp_true_on_policy.py', 'num_gpus': 2},
{'test_file': 'test_qwen3_vl_4B_fsdp.py', 'num_gpus': 8},
{'test_file': 'test_qwen3.5_0.8B_gsm8k_async_short.py', 'num_gpus': 4},
{'test_file': 'test_qwen3.5_0.8B_gsm8k_short.py', 'num_gpus': 4},
{'test_file': 'test_quick_start_glm4_9B.py', 'num_gpus': 8},
{'test_file': 'test_glm4.7_30B_A3B_pd_mooncake.py', 'num_gpus': 8},
{'test_file': 'test_qwen3_30B_A3B.py', 'num_gpus': 8},
{'test_file': 'test_qwen3.6_35B_A3B_pd_mooncake.py', 'num_gpus': 8, 'use_deepep': '1'},
{'test_file': 'test_qwen3_4B_ppo.py', 'num_gpus': 8},
{'test_file': 'test_moonlight_16B_A3B.py', 'num_gpus': 8},
{'test_file': 'test_mimo_7B_mtp_only_grad.py', 'num_gpus': 8},
{'test_file': 'test_qwen3_0.6B_parallel_check.py', 'num_gpus': 8},
{'test_file': 'test_qwen3_0.6B_megatron_fsdp_align.py', 'num_gpus': 4},
{'test_file': 'test_qwen3_4B_ckpt.py', 'num_gpus': 8},
{'test_file': 'test_qwen3_4B_ckpt.py --async-save', 'num_gpus': 8},
{'test_file': 'test_qwen3_4B_ckpt.py', 'test_args': '--async-save', 'num_gpus': 8},
{'test_file': 'test_qwen2.5_0.5B_debug_rollout_then_train.py', 'num_gpus': 8},
{'test_file': 'test_qwen2.5_0.5B_opd_sglang.py', 'num_gpus': 8},
],
Expand Down Expand Up @@ -109,24 +112,11 @@ jobs:
<% else %>
if: (github.event_name == 'workflow_dispatch') || (github.event.pull_request && contains(github.event.pull_request.labels.*.name, '<< config.label >>'))
<% endif %>
<% if config.get('cpu') %>
runs-on: ubuntu-latest
<% else %>
runs-on: self-hosted
container:
image: << config.image if config.image else 'slimerl/slime:latest' >>
options: >
--gpus all
--ipc=host
--shm-size=16g
--ulimit memlock=-1
--ulimit stack=67108864
--memory=0
--memory-swap=0
-e http_proxy=$http_proxy
-e https_proxy=$https_proxy
-e HTTP_PROXY=$HTTP_PROXY
-e HTTPS_PROXY=$HTTPS_PROXY
-v /mnt/nvme0n1/slime_ci:/data/slime_ci
-v /mnt/nvme0n1/slime_ci/models:/root/models
-v /mnt/nvme0n1/slime_ci/datasets:/root/datasets
<% endif %>
strategy:
fail-fast: false
matrix:
Expand All @@ -145,38 +135,101 @@ jobs:
steps:
- name: Checkout repository
uses: actions/checkout@v4
<% if config.get('cpu') %>

- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: '3.10'
cache: 'pip'

- name: Install dependencies
shell: bash
run: |
pip install torch --index-url https://download.pytorch.org/whl/cpu
pip install pytest numpy packaging pyyaml omegaconf tqdm httpx pybase64 pylatexenc sympy aiohttp pillow

- name: Install
shell: bash
run: cd $GITHUB_WORKSPACE && pip install -e . --no-deps --break-system-packages
run: cd $GITHUB_WORKSPACE && pip install -e . --no-deps
<% else %>
<% endif %>

- name: Execute
shell: bash
run: |
<% if config.get('cpu') %>
TEST_PATH="${{ matrix.info.test_file }}"
if [[ "$TEST_PATH" != tests/* ]]; then
TEST_PATH="tests/$TEST_PATH"
fi
TEST_ARGS="${{ matrix.info.test_args || '' }}"
if [[ -n "$TEST_ARGS" ]]; then
read -r -a TEST_ARGS_ARRAY < <(printf '%s\n' "$TEST_ARGS")
else
TEST_ARGS_ARRAY=()
fi
if [ "${{ matrix.info.num_gpus }}" = "0" ]; then
python "$TEST_PATH"
python "$TEST_PATH" "${TEST_ARGS_ARRAY[@]}"
else
python tests/ci/gpu_lock_exec.py --count ${{ matrix.info.num_gpus }} -- python "$TEST_PATH"
python tests/ci/gpu_lock_exec.py --count ${{ matrix.info.num_gpus }} -- python "$TEST_PATH" "${TEST_ARGS_ARRAY[@]}"
fi
<% else %>
docker run --rm \
--privileged \
--cap-add SYS_NICE \
--security-opt seccomp=unconfined \
--network host \
--gpus all \
--ipc=host \
--shm-size=16g \
--ulimit memlock=-1 \
--ulimit stack=67108864 \
--memory=0 \
--memory-swap=0 \
-e http_proxy \
-e https_proxy \
-e HTTP_PROXY \
-e HTTPS_PROXY \
-e GITHUB_COMMIT_NAME \
-e WANDB_API_KEY \
-e SLIME_TEST_ENABLE_INFINITE_RUN \
-e SLIME_TEST_USE_DEEPEP \
-e SLIME_TEST_USE_FP8_ROLLOUT \
-e SLIME_TEST_ENABLE_EVAL \
-e TEST_FILE="${{ matrix.info.test_file }}" \
-e TEST_ARGS="${{ matrix.info.test_args || '' }}" \
-e NUM_GPUS="${{ matrix.info.num_gpus }}" \
-v "$GITHUB_WORKSPACE:$GITHUB_WORKSPACE" \
-v /mnt/nvme0n1/slime_ci:/data/slime_ci \
-v /mnt/nvme0n1/slime_ci/models:/root/models \
-v /mnt/nvme0n1/slime_ci/datasets:/root/datasets \
-w "$GITHUB_WORKSPACE" \
<< config.image if config.image else 'slimerl/slime:latest' >> \
bash -lc '
set -euo pipefail
pip install -e . --no-deps --break-system-packages
TEST_PATH="$TEST_FILE"
if [[ "$TEST_PATH" != tests/* ]]; then
TEST_PATH="tests/$TEST_PATH"
fi
if [[ -n "$TEST_ARGS" ]]; then
read -r -a TEST_ARGS_ARRAY < <(printf "%s\n" "$TEST_ARGS")
else
TEST_ARGS_ARRAY=()
fi
if [ "$NUM_GPUS" = "0" ]; then
python "$TEST_PATH" "${TEST_ARGS_ARRAY[@]}"
else
python tests/ci/gpu_lock_exec.py --count "$NUM_GPUS" -- python "$TEST_PATH" "${TEST_ARGS_ARRAY[@]}"
fi
'
<% endif %>
<% endfor %>

e2e-test-changed-detect:
if: (github.event_name == 'workflow_dispatch') || (github.event.pull_request && contains(github.event.pull_request.labels.*.name, 'run-ci-changed'))
runs-on: self-hosted
container:
image: slimerl/slime:latest
options: >
--gpus all
--ipc=host
--shm-size=16g
--ulimit memlock=-1
--ulimit stack=67108864
--memory=0
--memory-swap=0
outputs:
matrix: ${{ steps.detect.outputs.matrix }}
has_tests: ${{ steps.detect.outputs.has_tests }}
Expand Down Expand Up @@ -217,23 +270,6 @@ jobs:
needs: e2e-test-changed-detect
if: needs.e2e-test-changed-detect.outputs.has_tests == 'true'
runs-on: self-hosted
container:
image: slimerl/slime:latest
options: >
--gpus all
--ipc=host
--shm-size=16g
--ulimit memlock=-1
--ulimit stack=67108864
--memory=0
--memory-swap=0
-e http_proxy=$http_proxy
-e https_proxy=$https_proxy
-e HTTP_PROXY=$HTTP_PROXY
-e HTTPS_PROXY=$HTTPS_PROXY
-v /mnt/nvme0n1/slime_ci:/data/slime_ci
-v /mnt/nvme0n1/slime_ci/models:/root/models
-v /mnt/nvme0n1/slime_ci/datasets:/root/datasets
strategy:
fail-fast: false
matrix: ${{ fromJson(needs.e2e-test-changed-detect.outputs.matrix) }}
Expand All @@ -252,19 +288,55 @@ jobs:
- name: Checkout repository
uses: actions/checkout@v4

- name: Install
shell: bash
run: cd $GITHUB_WORKSPACE && pip install -e . --no-deps --break-system-packages

- name: Execute
shell: bash
run: |
TEST_PATH="${{ matrix.info.test_file }}"
if [[ "$TEST_PATH" != tests/* ]]; then
TEST_PATH="tests/$TEST_PATH"
fi
if [ "${{ matrix.info.num_gpus }}" = "0" ]; then
python "$TEST_PATH"
else
python tests/ci/gpu_lock_exec.py --count ${{ matrix.info.num_gpus }} -- python "$TEST_PATH"
fi
docker run --rm \
--privileged \
--cap-add SYS_NICE \
--security-opt seccomp=unconfined \
--network host \
--gpus all \
--ipc=host \
--shm-size=16g \
--ulimit memlock=-1 \
--ulimit stack=67108864 \
--memory=0 \
--memory-swap=0 \
-e http_proxy \
-e https_proxy \
-e HTTP_PROXY \
-e HTTPS_PROXY \
-e GITHUB_COMMIT_NAME \
-e WANDB_API_KEY \
-e SLIME_TEST_ENABLE_INFINITE_RUN \
-e SLIME_TEST_USE_DEEPEP \
-e SLIME_TEST_USE_FP8_ROLLOUT \
-e SLIME_TEST_ENABLE_EVAL \
-e TEST_FILE="${{ matrix.info.test_file }}" \
-e TEST_ARGS="${{ matrix.info.test_args || '' }}" \
-e NUM_GPUS="${{ matrix.info.num_gpus }}" \
-v "$GITHUB_WORKSPACE:$GITHUB_WORKSPACE" \
-v /mnt/nvme0n1/slime_ci:/data/slime_ci \
-v /mnt/nvme0n1/slime_ci/models:/root/models \
-v /mnt/nvme0n1/slime_ci/datasets:/root/datasets \
-w "$GITHUB_WORKSPACE" \
slimerl/slime:latest \
bash -lc '
set -euo pipefail
pip install -e . --no-deps --break-system-packages
TEST_PATH="$TEST_FILE"
if [[ "$TEST_PATH" != tests/* ]]; then
TEST_PATH="tests/$TEST_PATH"
fi
if [[ -n "$TEST_ARGS" ]]; then
read -r -a TEST_ARGS_ARRAY < <(printf "%s\n" "$TEST_ARGS")
else
TEST_ARGS_ARRAY=()
fi
if [ "$NUM_GPUS" = "0" ]; then
python "$TEST_PATH" "${TEST_ARGS_ARRAY[@]}"
else
python tests/ci/gpu_lock_exec.py --count "$NUM_GPUS" -- python "$TEST_PATH" "${TEST_ARGS_ARRAY[@]}"
fi
'
2 changes: 1 addition & 1 deletion CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ Thank you for your interest in contributing to slime! We deeply appreciate every

## Collaboration Scope

slime is the RL training infrastructure behind [GLM-4.5 through GLM-5](https://z.ai) and a large number of internal experiments at Z.ai. We open-sourced slime because we believe the training scenarios used internally cover the majority of cutting-edge RL algorithm requirements, and we hope to provide the community with a correct and efficient large-scale RL training infrastructure.
slime is the RL training infrastructure behind [GLM-4.5 through GLM-5.1](https://z.ai) and a large number of internal experiments at Z.ai. We open-sourced slime because we believe the training scenarios used internally cover the majority of cutting-edge RL algorithm requirements, and we hope to provide the community with a correct and efficient large-scale RL training infrastructure.

Our goal for open-source collaboration is focused on **bug fixes** and **general-purpose large-scale RL optimizations**. We have had several successful collaborations with the community in this area, including:

Expand Down
12 changes: 10 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,8 +10,8 @@
1. **High-Performance Training**: Supports efficient training in various modes by connecting Megatron with SGLang;
2. **Flexible Data Generation**: Enables arbitrary training data generation workflows through custom data generation interfaces and server-based engines.

slime is the RL-framework behind [GLM-5](https://z.ai/blog/glm-5), [GLM-4.7](https://z.ai/blog/glm-4.7), [GLM-4.6](https://z.ai/blog/glm-4.6), [GLM-4.5](https://z.ai/blog/glm-4.5) and apart from models from Z.ai, we also supports the following models:
- Qwen3 series (Qwen3Next, Qwen3MoE, Qwen3), Qwen2.5 series;
slime is the RL-framework behind [GLM-5.1](https://z.ai/blog/glm-5.1), [GLM-5](https://z.ai/blog/glm-5), [GLM-4.7](https://z.ai/blog/glm-4.7), [GLM-4.6](https://z.ai/blog/glm-4.6), [GLM-4.5](https://z.ai/blog/glm-4.5) and apart from models from Z.ai, we also supports the following models:
- Qwen series (Qwen3.6, Qwen3.5, Qwen3Next, Qwen3MoE, Qwen3, Qwen2.5);
- DeepSeek V3 series (DeepSeek V3, V3.1, DeepSeek R1);
- Llama 3.

Expand Down Expand Up @@ -51,6 +51,14 @@ We also provide examples for some use cases not covered in the quick start guide

slime has powered several novel research projects and production systems. Here are some notable examples:

### 🌈 Relax: Asynchronous RL Engine for Omni-Modal Agentic Training

[**Relax**](https://github.com/redai-infra/Relax) (Reinforcement Engine Leveraging Agentic X-modality) is an omni-modal agentic RL framework open-sourced by the RedAI Infra team, built upon the slime infrastructure stack that combines Ray, Megatron-LM, and SGLang. Relax adopts a service-oriented architecture on Ray Serve with Megatron-LM and SGLang as training/inference backends. It uses [TransferQueue](https://github.com/Ascend/TransferQueue) to fully decouple Actor, Rollout, ActorFwd, Reference, and Advantage computation onto independent GPU clusters, and introduces **DCS (Distributed Checkpoint Service)** — an NCCL-broadcast weight-sync engine that streams updated Actor weights to Rollout/ActorFwd/Reference asynchronously and overlaps the transfer with the next training step, enabling fully-async training at configurable staleness. Relax supports end-to-end RL for text, vision, and audio (including Qwen3-Omni) and agentic multi-turn rollouts.

### 🦞 OpenClaw-RL: Train a Personalized Clawbot Simply by Talking to It

[**OpenClaw-RL**](https://github.com/Gen-Verse/OpenClaw-RL) is an RL server for personalized OpenClaw agents. It hosts the OpenClaw model and improves it from prior conversations across deployments, while slime's asynchronous RL infrastructure prevents training from interfering with API serving. It supports two automatic optimization methods: GRPO with binary feedback inferred from subsequent states, and on-policy distillation that extracts hindsight hints from later feedback for the current policy.

### ⚛️ P1: Mastering Physics Olympiads with Reinforcement Learning

[**P1**](https://prime-rl.github.io/P1/) is a family of open-source physics reasoning models trained entirely through reinforcement learning. P1 leverages slime as the RL post training framework, and introduces a multi-stage RL training algorithm that progressively enhances reasoning ability through adaptive learnability adjustment and stabilization mechanisms. Enpowered by this training paradigm, P1 delivers breakthrough performance in open-source physics reasoning.
Expand Down
4 changes: 2 additions & 2 deletions README_zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,8 +10,8 @@
1. **高性能训练**:通过连接 Megatron 与 SGLang,支持各种模式的高效训练;
2. **灵活的数据生成**:通过自定义数据生成接口以及 server based engine,实现任意的数据训练数据生成流程。

slime 是 [GLM-5](https://z.ai/blog/glm-5)、[GLM-4.7](https://z.ai/blog/glm-4.7)、[GLM-4.6](https://z.ai/blog/glm-4.6)、[GLM-4.5](https://z.ai/blog/glm-4.5) 背后的 RL 训练框架,除此之外,slime 还支持:
- Qwen3 系列 (Qwen3Next, Qwen3MoE, Qwen3), Qwen2.5 系列
slime 是 [GLM-5.1](https://z.ai/blog/glm-5.1)、[GLM-5](https://z.ai/blog/glm-5)、[GLM-4.7](https://z.ai/blog/glm-4.7)、[GLM-4.6](https://z.ai/blog/glm-4.6)、[GLM-4.5](https://z.ai/blog/glm-4.5) 背后的 RL 训练框架,除此之外,slime 还支持:
- Qwen 系列 (Qwen3.6、Qwen3.5、Qwen3NextQwen3MoEQwen3Qwen2.5)
- DeepSeek V3 系列 (DeepSeek V3, V3.1, DeepSeek R1);
- Llama 3。

Expand Down
Loading
Loading