Skip to content

Add support for NVIDIA DGX Spark (GB10 / sm_121a, arm64) - #1835

Merged
zhuzilin merged 7 commits into
THUDM:mainfrom
boots-coder:feat/gb10-support
Apr 21, 2026
Merged

Add support for NVIDIA DGX Spark (GB10 / sm_121a, arm64)#1835
zhuzilin merged 7 commits into
THUDM:mainfrom
boots-coder:feat/gb10-support

Conversation

@boots-coder

Copy link
Copy Markdown
Contributor

Add NVIDIA DGX Spark (GB10 / sm_121a) support

Summary

This PR adds support for NVIDIA DGX Spark (Project Digits, GB10 chip — Grace
CPU + consumer Blackwell GPU at sm_121a, aarch64, 128 GB unified memory) to
slime's docker build pipeline. GB10 is an explicit gap in the current support
matrix:

  • slime's published images (slimerl/slime:*) are x86_64-only.
  • The arm64 base slimerl/sglang:v0.5.9 ships CUDA 12.9 whose ptxas has no
    sm_121a target, so triton JIT crashes on first kernel.
  • The ENABLE_CUDA_13=1 branch in the upstream Dockerfile is aimed at GB200/GB300
    (sm_100a) with an x86-only sgl-router wheel.

This PR rebases a second Dockerfile (docker/Dockerfile.gb10) on
nvcr.io/nvidia/vllm:26.03-py3 (arm64), which ships CUDA 13.2, PyTorch 2.11.0a0
with compute_120 PTX, Triton 3.6.0, and flash-attn 2.7.4.post1 — the minimal
viable baseline for GB10. Fifteen small blockers are resolved in the Dockerfile
and supporting patches.

End-to-end validation: Qwen2.5-0.5B + GRPO + dapo-math-17k, 1 GB10 GPU,
colocated actor + rollout, one full rollout → reward → policy-update cycle
completed in 2m10s (step 0 metrics logged, weight sync to SGLang succeeded,
checkpoint saved).

What's in the patch

New files

  • docker/Dockerfile.gb10 — digest-pinned, arm64-native build on NGC vllm base
  • docker/patch/gb10/patch_sgl_kernel.py — adds SGL_KERNEL_GB10_ONLY CMake option
  • docker/patch/gb10/sgl-kernel-arch.patch — unified diff of the same
  • docker/patch/gb10/cuda_profiler_api.h — 20-line shim for a CUDA 13 dropped header
  • scripts/run-qwen2.5-0.5B-gb10-smoke.sh — minimal 1-GPU smoke script
  • NOTES_GB10.md — walkthrough of the 15 blockers with root causes

Existing file edits

  • slime/utils/arguments.py — remove trailing comma that turned an argparse
    help= string into a single-element tuple (causes 'tuple' has no attribute 'strip'
    during --help; affects all platforms, not just GB10)

The 15 blockers (quick reference)

# Blocker Resolution
1 ptxas fatal: sm_121a Rebase on NGC CUDA 13.2 image
2 libnvrtc.so.12 missing from sgl_kernel wheel Build sgl_kernel from source
3 sgl_kernel wheel ABI-mismatch vs NGC libtorch Build against NGC torch
4 sgl_kernel arm64 wheels only have sm_100 variant Build sm_121 variant
5 sgl_kernel source build OOM (7 archs × cutlass) SGL_KERNEL_GB10_ONLY CMake option
6 TE: CUDNN::cudnn_engines_precompiled not found Install nvidia-cudnn-cu13==9.20.0.48 pypi wheel, symlink
7 TE: nvtx3/nvToolsExt.h missing Copy headers from NVIDIA/NVTX repo
8 TE ptx.cuh: "use smXXXf, not smXXX" NVTE_CUDA_ARCHS="120f;121f"
9 CMake 3.31 rejects f suffix Upgrade to cmake==4.3.1, override PIP_CONSTRAINT
10 cuda_profiler_api.h missing in CUDA 13 20-line shim header
11 'tuple' object has no attribute 'strip' Remove trailing comma in arguments.py
12 sglang_router x86_64 only Use upstream sglang-router==0.3.2 (arm64 wheel)
13 antlr4==4.13.2 breaks omegaconf Pin antlr4-python3-runtime==4.9.3
14 megatron.training not found post pip install -e PYTHONPATH=/root/src/Megatron-LM
15 tilelang needs libz3.so apt install libz3-dev

Testing

Run inside the image:

cd /root/slime
python train.py --help   # prints 3714 lines, exits 0
bash scripts/run-qwen2.5-0.5B-gb10-smoke.sh   # end-to-end GRPO run, 2m10s

Scope this PR does NOT cover

  • FA3 (Hopper-only) kernels: deliberately skipped; GB10 cannot run sm_90a binaries.
  • Any multi-GPU or multi-node GB10 topology (DGX Spark is a single-GPU SKU).
  • Convergence / training quality validation (separate ongoing benchmark).

Reproducibility

cd slime
docker build -f docker/Dockerfile.gb10 -t slime-gb10:latest .

Base image is pinned by digest:

nvcr.io/nvidia/vllm@sha256:13e327dad79e6e417f6687fec2ba76b0386d597082ec0ee003c1e964ec6ad0e7

@boots-coder

Copy link
Copy Markdown
Contributor Author
image image image End-to-end smoke on a real GB10 (DGX Spark). Qwen2.5-0.5B + GRPO + dapo-math-17k, 1 GPU colocated, 1 rollout × 1 policy step. Full cycle: 1m26s, Ray job succeeded, 8/8 weight-sync HTTP calls returned 200 OK.

@boots-coder
boots-coder force-pushed the feat/gb10-support branch 2 times, most recently from 37561d6 to 9447a85 Compare April 15, 2026 08:10
The help= kwarg was written as a single-element tuple (trailing comma),
which makes argparse format_help() raise:
  AttributeError: 'tuple' object has no attribute 'strip'
on any `python train.py --help` invocation.

Remove the comma so the help text is a plain string.
slime's published images are x86_64-only. The arm64 slimerl/sglang base
ships CUDA 12.9 whose ptxas lacks an sm_121a target, so Triton crashes on
first JIT. The ENABLE_CUDA_13=1 branch in the upstream Dockerfile is
aimed at GB200/GB300 (sm_100a) with an x86-only sgl-router wheel and does
not work on GB10.

This change adds a second Dockerfile targeting GB10 specifically:

  - Rebased on nvcr.io/nvidia/vllm:26.03-py3 (arm64), pinned by digest.
    Ships CUDA 13.2, PyTorch 2.11.0a0 with compute_120 PTX,
    Triton 3.6.0, flash-attn 2.7.4.post1.

  - sgl-kernel is rebuilt from source with a new CMake option
    SGL_KERNEL_GB10_ONLY=ON (patch_sgl_kernel.py / sgl-kernel-arch.patch)
    that restricts gencode to sm_120a + sm_121a. The stock 7-arch
    emission OOM-kills cicc on 128 GB Spark hosts (each extra cutlass
    FP8 gemm arch costs ~10-15 GB RAM per TU).

  - TransformerEngine 2.10 is built with NVTE_CUDA_ARCHS=120f;121f.
    The 'f' (family-specific) arch suffix is required by TE's ptx.cuh
    static_assert and is only parsed by CMake >= 4.0, so cmake is
    upgraded to 4.3.1 (NGC's PIP_CONSTRAINT must be cleared).

  - Small CUDA 13 gaps that NGC does not ship are filled:
    * cuda_profiler_api.h shim (symbols remain in libcudart.so)
    * NVTX3 headers (copied from github.com/NVIDIA/NVTX)
    * libcudnn_engines_precompiled.so.9 (from nvidia-cudnn-cu13 wheel)

  - The x86-only zhuzilin/sgl-router wheel is swapped for upstream
    sglang-router==0.3.2 (has arm64 wheel). slime's version-compare
    code still works; only the 'slime' in version wandb branch differs.

Also adds scripts/run-qwen2.5-0.5B-gb10-smoke.sh: a 1-GPU colocated
smoke that exercises the full rollout -> reward -> policy-update ->
weight-sync cycle. Validated end to end with Qwen2.5-0.5B on
dapo-math-17k; a single step completes in ~2m10s on GB10.
Captures the root causes and resolutions for every issue encountered
while porting the full stack (sgl-kernel, TransformerEngine, apex,
Megatron-LM, sglang, slime) to GB10 / sm_121a. Placed next to
Dockerfile.gb10 so reviewers see the reference material when reading
the Dockerfile.
…nc, math_verify, pyarrow, accelerate)

These packages were installed manually during the interactive GB10 build
session but were missing from the Dockerfile, breaking clean reproduction.
- numpy<2: Megatron requires numpy 1.x (NGC ships 2.x)
- pylatexenc, math_verify, word2number: reward function runtime deps
- pyarrow: GSM8K parquet data preprocessing
- accelerate: HuggingFace transformers device_map support
@boots-coder

boots-coder commented Apr 16, 2026

Copy link
Copy Markdown
Contributor Author

100-rollout convergence run + held-out test set evaluation on GB10

Following the smoke test posted earlier, I ran a full 100-rollout GRPO convergence experiment on the GB10 (single GPU, colocated mode), then evaluated on the held-out GSM8K test set (1,319 questions, greedy decoding).

Test Set Results (Pass@1, greedy, 1319 questions)

Model Correct Accuracy No \boxed{} output Avg response length
Baseline (Qwen2.5-0.5B-Instruct) 575 / 1319 43.59% 33 (2.5%) 1071 chars
+ 100-step GRPO (iter_99) 662 / 1319 50.19% 3 (0.2%) 907 chars
Improvement +87 +6.60 pp −30 (format compliance ↑) −15.3% (more concise)

Three dimensions improved simultaneously:

  1. Accuracy +6.60 pp on never-seen test questions
  2. Format compliance 97.5% → 99.8% — model learned to reliably output \boxed{}
  3. Responses 15% shorter — RL optimized for quality, not length

Training Configuration

  • Model: Qwen2.5-0.5B-Instruct (494M params)
  • Data: GSM8K train split (7,473 questions), --rm-type math (grade_answer_verl)
  • Scale: 20 prompts × 8 samples × 100 rollouts = 16,000 total samples (< 1 epoch)
  • Hyperparams: lr=1e-6 constant, GRPO with eps-clip=0.2/0.28, colocated single GPU
  • Training time: 1h 11min (41.9s/rollout avg)
  • Eval time: 28min (2 × 1319 greedy, HF transformers backend)

Training Stability

Metric Value
grad_norm 0.93 mean (range 0.73–1.18, no explosion)
entropy_loss 0.295 → 0.172 (monotonic decrease)
kl_loss vs ref 0.0 → 0.044 (smooth, controlled)
Rollout throughput 5,495 tok/gpu/s
Actor train TFLOPS 16.83 (mean)

Reward Curve (training rollout, temp=1.0)

fig1_reward_curve fig2_stability fig3_length

Rollout raw_reward trajectory: 0.375 → 0.656 (peak 0.763 at rollout 58).
Note: rollout metrics use temp=1.0 sampling on training data (8 samples/prompt), so they're higher than the greedy test-set Pass@1 reported above.

Reproduction

All scripts, configs, and the updated Dockerfile.gb10 (with 6 missing runtime deps fixed in commit 8f55ecb1) are in this PR.

@zhuzilin

Copy link
Copy Markdown
Contributor

This is amazing! Thank you!

@zhuzilin
zhuzilin merged commit f44af13 into THUDM:main Apr 21, 2026
SamitHuang added a commit to SamitHuang/slime that referenced this pull request May 17, 2026
* temp save rfc

Signed-off-by: SamitHuang <285365963@qq.com>

* add plan

Signed-off-by: SamitHuang <285365963@qq.com>

* update

Signed-off-by: SamitHuang <285365963@qq.com>

* [docker] remove true on policy patches (THUDM#1661)

Co-authored-by: Copilot <copilot@github.com>

* [fix]: Qwen3.5-35B-A3B 8-GPU: set TP size to 2 for num_query_groups=2 (THUDM#1662)

* Remove FSDP support (THUDM#1664)

Co-authored-by: Copilot <copilot@github.com>

* docs: add OpenClaw-RL to projects built upon slime (THUDM#1635)

* qwen2.5 0.5b non-colocate (first attempt ok, but nccl error later)

Signed-off-by: samithuang <285365963@qq.com>

* add convert script

* add setup doc

* Support setting update weights in sglang_config (THUDM#1665)

Co-authored-by: Copilot <copilot@github.com>

* fix nccl error by NcclBridge subprocess

* eliminate gpu to cpu weight transfer

Signed-off-by: samithuang <285365963@qq.com>

* Revise weight synchronization strategy in goal plan

Reorder weight synchronization support for colocate and non-colocate scenarios in the goal plan.

* [fix] Fix numerical accuracy issue in dynamic sampling filter (THUDM#1674)

* sync from internal (THUDM#1677)

Co-authored-by: Copilot <copilot@github.com>

* bugfixes from community (THUDM#1678)

Co-authored-by: Copilot <copilot@github.com>
Co-authored-by: yueming-yuan <yym022502@gmail.com>
Co-authored-by: coding-famer <chenhegu0109@gmail.com>

* Fix: pass return_tensors in text_kwargs for transformers>=5.0.0 compatibility (THUDM#1648)

* Fix missing packed_seq_params in bshd qkv_format (THUDM#1649)

* [Multimodal][Model] Qwen3.5 VL training example/support (THUDM#1676)

* update docs (THUDM#1680)

Co-authored-by: Copilot <copilot@github.com>

* update docs (THUDM#1681)

Co-authored-by: Copilot <copilot@github.com>

* support offloading non-updatable server (THUDM#1668)

Co-authored-by: Copilot <copilot@github.com>

* bugfix (THUDM#1685)

Co-authored-by: Copilot <copilot@github.com>

* fix: handle Qwen3.5 in quantize_params_fp8 (THUDM#1683)

* bugfix (THUDM#1687)

Co-authored-by: Copilot <copilot@github.com>

* Fix Qwen3.5 & Qwen3-Next linear attention cu_seqlens missing (THUDM#1686)

Co-authored-by: benyi <huangliangmeng.hlm@alibaba-inc.com>

* fix: use semantic version comparison for PyTorch >= 2.6 detection (THUDM#1667)

* [Fix] Minor fix for properly finishing / flushing wandb logging metrics at exit (THUDM#1592)

Co-authored-by: Zilin Zhu <zhuzilinallen@gmail.com>

* Autofix/issue 1578 hf2megatron arg suffix (THUDM#1636)

* bugfix (THUDM#1688)

Co-authored-by: Copilot <copilot@github.com>

* fix(examples): update strands_sglang example to v0.3.x API (THUDM#1684)

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>

* [docker] cherry pick qwen3.5 bugfix (THUDM#1691)

Co-authored-by: Copilot <copilot@github.com>

* bugfix/fix Qwen3.5 dense model precision bug in TP_SIZE>1 from sglang (THUDM#1705)

* Fix/qwen3 5 mtp bridge (THUDM#1702)

Co-authored-by: benyi <huangliangmeng.hlm@alibaba-inc.com>

* support epd for glm4.6v (THUDM#1704)

* [docker] support epd for glm4.6v (THUDM#1707)

Co-authored-by: Copilot <copilot@github.com>

* remove script

* [docker] store v0.5.9 patch (THUDM#1710)

Co-authored-by: Copilot <copilot@github.com>

* Add GLM-4.7-Flash MTP training support (THUDM#1712)

* [release] bump to v0.2.3 (THUDM#1682)

Co-authored-by: Copilot <copilot@github.com>

* feat: add GLM-4.6V MoE VL bridge with CP support (THUDM#1715)

Co-authored-by: Copilot <copilot@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* fix: resolve rope_theta from rope_parameters dict in HF config validation (THUDM#1720)

Co-authored-by: Copilot <copilot@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* [docker] patches for glm4.6v, kimi k2.5 and dsa cp only (THUDM#1722)

Co-authored-by: Copilot <copilot@github.com>

* [docker] support IndexCache

* Fix CUDA IPC cache leaks during weight updates (THUDM#1731)

Co-authored-by: Copilot <copilot@github.com>

* [docker] update megatron (THUDM#1729)

Co-authored-by: Copilot <copilot@github.com>

* [docker] Fix IndexCache with mla model (THUDM#1736)

Co-authored-by: Copilot <copilot@github.com>

* [slime-router] support pd disaggregation and remove radix tree middleware (THUDM#1735)

* Fix glm4v megatron bridge (THUDM#1738)

Co-authored-by: Copilot <copilot@github.com>

* [docker] update sglang patch (THUDM#1743)

Co-authored-by: Copilot <copilot@github.com>

* feat: GLM4V multimodal support improvements (THUDM#1745)

Co-authored-by: Copilot <copilot@github.com>

* feat: placeholder worker type, metrics router, and GPQA letter range (THUDM#1746)

Co-authored-by: Copilot <copilot@github.com>

* always enable_metrics and remove dp context (THUDM#1747)

Co-authored-by: Copilot <copilot@github.com>

* fix: resolve SP/CP gradient inflation in FLA (linear attention) layers (THUDM#1748)

Co-authored-by: Copilot <copilot@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Update MTP example configs, rename GLM-4.5 to GLM-4.7, clean scripts (THUDM#1749)

Co-authored-by: Copilot <copilot@github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Support qwen3.5 loss mask for multi-turn SFT (THUDM#1742)

Co-authored-by: benyi <huangliangmeng.hlm@alibaba-inc.com>

* fix: propagate moe_token_dispatcher_type in bridge model provider (THUDM#1737)

* fix: resolve rope_theta from rope_parameters in DeepseekV32Bridge (THUDM#1734)

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>

* chore: translate remaining Chinese comments to English (THUDM#1726)

* feat: add Qwen3.5-4B model support (THUDM#1721)

* fix: http_utils. disable system proxy for internal SGLang httpx clients (THUDM#1714)

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: auto-detect GPUs in qwen3-4b script (THUDM#1700)

* fix: quote `$MOE_LAYER_FREQ` (THUDM#1689)

* disable router health_check and allow prompt_data is None (THUDM#1751)

Co-authored-by: Copilot <copilot@github.com>

* Router for vllm (#5)

* Draft router design

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Add vllm router

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Add router to script

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix gpu memory utilization

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix output token ids

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Add more nccl flag

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix bug

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

---------

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* small fix on qwen3-235b-a22b launch script (THUDM#1719)

* sync internal bugfix (THUDM#1765)

Co-authored-by: Copilot <copilot@github.com>

* Fix uploading sglang metrics to wandb (THUDM#1768)

Co-authored-by: Copilot <copilot@github.com>

* use zhuzilin/sgl-router for sglang-router (THUDM#1770)

Co-authored-by: Copilot <copilot@github.com>

* [docker] update sgl-router (THUDM#1772)

Co-authored-by: Copilot <copilot@github.com>

* [Multimodal] Add Multimodal OPD support (THUDM#1760)

* refactor: remove slime router (THUDM#1773)

Co-authored-by: Copilot <copilot@github.com>

* Add rollout trace timeline viewer (THUDM#1776)

Co-authored-by: Hanyu Zhang <hanyu.zhang@aminer.cn>

* [Fix] Fix duplicate Megatron LR scheduler resume when optimizer state is not loaded (THUDM#1775)

* Support FP8 conversion for Qwen3.5 (THUDM#1769)

* fix typo (THUDM#1759)

Co-authored-by: shiqirui <shiqirui@kupasai.com>

* [Fix]Fix some bugs/clean up (THUDM#1756)

* (fix):not have encoder_only attr cause run failed (THUDM#1741)

Co-authored-by: wangch <wangch@wangchdeMacBook-Air.local>

* update docs

* remove redundant envvar

* some minor cleanup

* [release] bump to v0.2.4  (THUDM#1777)

Co-authored-by: Copilot <copilot@github.com>

* Plan refactor vllm/sglang

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Code implemented

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix bug

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix bug

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix bug

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix port

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix config

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix bug MOE weight sync

* Fix bug vllm transfer weight

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix weight sync

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix config

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Change name config

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* pass critic role through to create RayTrainGroup (THUDM#1797)

* fix qwen3.5 397B converting error when enable expert parallel (THUDM#1799)

Co-authored-by: 周鹤云 <zhouheyun@xiaohongshu.com>

* fix(geo3k-vlm-sft): remove --apply-chat-template from SFT launch script (THUDM#1791)

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>

* Add host memory metrics to available_memory function (THUDM#1764)

* [WIP] fix loss oom (THUDM#1788)

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* sync from internal (THUDM#1805)

* sync from internal (THUDM#1807)

* feat: add npu patch for qwen3-vl-8b grpo & ppo (#1750)

Signed-off-by: cjy0x <isjunyi.chen@gmail.com>
Co-authored-by: shiyuan680 <917935075@qq.com>
Co-authored-by: PengchengShi00 <spc117369@gmail.com>

* fix missing position_ids in log-prob forward step (THUDM#1809)

* feat: add support for including missing weights from origin HF checkp… (THUDM#1812)

* [Fix] Initialize grad_norm before found_inf skip path (THUDM#1762)

* [conda] Add install custom sgl-router to build_conda.sh (THUDM#1813)

* Revert no_grad for entropy to prevent comm stuck in dsa (THUDM#1822)

* Add fallback for get_seqlen_balanced_partitions (THUDM#1823)

* Resolve review

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Try colocated vllm weight

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* docs: add Relax to notable projects in README (THUDM#1834)

* Bugfix: use cpu instead of cuda in convert_torch_dist_to_hf.py when --add-missing-from-origin-hf is set (THUDM#1828)

* [fix] eval sample logging when sample is a list (THUDM#1836)

* [Draft] Local runable dev

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* [Fix]  Fix cuda-python pin in build_conda.sh (THUDM#1827)

* fix entropy bug and update code (THUDM#1846)

* Revert "Add fallback for get_seqlen_balanced_partitions" (THUDM#1848)

* fix (THUDM#1849)

* Fix offload train

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Add support for NVIDIA DGX Spark (GB10 / sm_121a, arm64) (THUDM#1835)

* Fix offload train

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix offload_rollout

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix vllm offload

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix offload traing

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix offload weight

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix offload weight

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* refactor/ppo (THUDM#1856)

* [docker] cleanup sglang patch (THUDM#1859)

* [docker] update v0.5.9 patch

* Rename critic config to megatron config (THUDM#1866)

* [Fix] Use Ray ObjectRef await instead of asyncio.to_thread in distributed POST (THUDM#1873)

* chore: include length context in slice_log_prob_with_cp assert (THUDM#1862)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* [docker] upgrade megatron to 1dcf0dafa (THUDM#1867)

* fix ppo value head load bugs (THUDM#1878)

* [docker] upgrade sglang to v0.5.10.post1 (THUDM#1874)

* [docs] update docs

* [docker] update megatron-bridge and add qwen3.6 tests (THUDM#1884)

* fix lint

* Fix(checkpoint): add resume/pause in save_model() for offload_train (fixes THUDM#1886) (THUDM#1888)

* fix ppo value offload bugs (THUDM#1882)

* fix qwen3.6 hf config validation bug (THUDM#1889)

* Add missing metrics to log (THUDM#1890)

* fix(qwen3_next): use torch.get_default_dtype() — get_current_dtype do… (THUDM#1883)

Co-authored-by: yeqinghe <yeqinghe@MacBook-Pro-6.local>

* Fix location error in install script (THUDM#1877)

* Only allow --allgather-cp for DSA model (THUDM#1891)

* Migrate internal feature (THUDM#1897)

* [Fix]  Fix distributed POST actor concurrency split (THUDM#1880)

Co-authored-by: Zilin Zhu <zhuzilinallen@gmail.com>

* Fix CI: update rollout_data_postprocess plugin contract for new call site (THUDM#1902)

Co-authored-by: jingshenghang <shenghang.jing@aminer.cn>

* Patch Megatron TP grad coalesce to chunked all-reduce (THUDM#1899)

* fix: harden retool rollout against multi-turn / retry desync (THUDM#1861)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Fix log file

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix import engine group

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

* Fix rebase code

Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>

---------

Signed-off-by: SamitHuang <285365963@qq.com>
Signed-off-by: samithuang <285365963@qq.com>
Signed-off-by: knlnguyen1802 <knlnguyen1802@gmail.com>
Signed-off-by: cjy0x <isjunyi.chen@gmail.com>
Co-authored-by: SamitHuang <285365963@qq.com>
Co-authored-by: Zilin Zhu <zhuzilinallen@gmail.com>
Co-authored-by: Copilot <copilot@github.com>
Co-authored-by: none0663 <none0663@outlook.com>
Co-authored-by: Yinjie Wang <yinjie@uchicago.edu>
Co-authored-by: Fengqing Jiang <43953876+Django-Jiang@users.noreply.github.com>
Co-authored-by: yueming-yuan <yym022502@gmail.com>
Co-authored-by: coding-famer <chenhegu0109@gmail.com>
Co-authored-by: Lawrence Wu <lawrence.wu@harmonic.fun>
Co-authored-by: huang3eng <huang3eng@gmail.com>
Co-authored-by: benyi <huangliangmeng.hlm@alibaba-inc.com>
Co-authored-by: Aaron Batilo <AaronBatilo@gmail.com>
Co-authored-by: Silun Wang <igeekwang@gmail.com>
Co-authored-by: Chengxing Xie <91449279+yitianlian@users.noreply.github.com>
Co-authored-by: Yuan He <33579950+Lawhy@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Mor Zusman <mor.zusmann@gmail.com>
Co-authored-by: append-only <shw20010329@163.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Nan Jiang <59716405+nanjiangwill@users.noreply.github.com>
Co-authored-by: Xuan Wang <49010704+stevewx@users.noreply.github.com>
Co-authored-by: Hubert Wang <huberthyw@gmail.com>
Co-authored-by: Hou Shihao <shhou007@gmail.com>
Co-authored-by: DongzhuoranZhou <110855293+DongzhuoranZhou@users.noreply.github.com>
Co-authored-by: Ailuntz <130897222+ailuntz@users.noreply.github.com>
Co-authored-by: Zhuohao Li <garrick0508@gmail.com>
Co-authored-by: Hanyu Zhang <hanyu.zhang@aminer.cn>
Co-authored-by: Kang Yu <kangy.me@gmail.com>
Co-authored-by: peterjc123 <peter_jiachen@163.com>
Co-authored-by: qrskannbara <94727257+albaNnaksqr@users.noreply.github.com>
Co-authored-by: shiqirui <shiqirui@kupasai.com>
Co-authored-by: wangyufak <wangch9@xiaopeng.com>
Co-authored-by: wangch <wangch@wangchdeMacBook-Air.local>
Co-authored-by: Xintong Li <znculee@gmail.com>
Co-authored-by: TM <tianmingxu.tmxu@gmail.com>
Co-authored-by: 周鹤云 <zhouheyun@xiaohongshu.com>
Co-authored-by: LiLei <77353389+lilei199908@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: cjy0x <isjunyi.chen@gmail.com>
Co-authored-by: shiyuan680 <917935075@qq.com>
Co-authored-by: PengchengShi00 <spc117369@gmail.com>
Co-authored-by: 杨睿 <595403043@qq.com>
Co-authored-by: Mathew Han <49226490+mathewjhan@users.noreply.github.com>
Co-authored-by: haoxuanJIA <116806014+boots-coder@users.noreply.github.com>
Co-authored-by: ryang <38470282+ryang-max@users.noreply.github.com>
Co-authored-by: Leo Fan <84952531+leofan-lab@users.noreply.github.com>
Co-authored-by: Long Yijun <156500868+Procrastinatorrrr@users.noreply.github.com>
Co-authored-by: HeatherLiuzh <heather996lzh@gmail.com>
Co-authored-by: yeqinghe <yeqinghe@MacBook-Pro-6.local>
Co-authored-by: tao W <122036357+selfanti@users.noreply.github.com>
Co-authored-by: jingshenghang <48083555+jingshenghang@users.noreply.github.com>
Co-authored-by: jingshenghang <shenghang.jing@aminer.cn>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants