Multi-Stage Awake: Support tag-based Resume and Pause #20

hebiao064 · 2025-06-11T06:03:54Z

Motivation

In RL Ecosystem which use colocate design like verl, we need to offload training model and load serving model & KV Cache frequently.

Background

Currently SGLang is using torch_memory_saver to pause and resume.
torch_memory_saver is a open source repo that provided easy to use api to hack cudaMalloc and cudaFree to make sure the virtual address could be consistent after pause and resume, which is critical to ensure CUDA Graph work.
CUDA Graph is critical to make sure SGLang runs faster in decoding phases.

Here is the current behavior of VERL + SGLang

During Training, we have training model and optimizer state in the GPU Memory, and once training is done, we will offload optimizer state to cpu and keep the model weights in GPU, which is needed in Update Weight.
During Update Weight, we awake the SGLang engine, so those paused memory of Model Weights and KV Cache will come back. Then we update model from training model to serving model on the fly using the api: update_weights_in_tensor
After Model being updated, we delete the training model from GPU Memory.

Above design works pretty well so far, however, this would waste a big chunk of GPU Memory during rollout, which could cause a few issues we've seen so far:

Small KV Cache: We need to use relative lower number of mem fraction ratio (e.g: 0.6), hence our KV Cache has less tokens. Given KV Cache has less tokens, we will hit RuntimeError: Prefill out of memory. Try to lower your batch size. when we try prefill large number of requests.
Out of Memory: If we use mem fraction ratio 0.8 and run RL for 32B model on 8 H100, it will OOM during update weight

Proposal

During Training, we do the same
During Update Weight Stage 1, we awake the model weights from SGLang and then update weights
During Update Weight Stage 2, we delete the training model weights from GPU Memory
Awake the SGLang's KV Cache

Benefit

With above feature, we can train larger model with same GPU, we can also make training/rollout more efficient given we can allocate larger KV Cache

Execution Plan: Keep using Singleton and provide tag based pause/resume

Support tag based resume/pause: Multi-Stage Awake: Support tag-based Resume and Pause #20
Support Multiple Stage Awake in SGLang: Multi-Stage Awake: Support Resume and Pause KV Cache and Weights separately sgl-project/sglang#7099
Support Multiple Stage Awake in verl: [rollout] feat: Support Multi-stage Awake for SGLang volcengine/verl#1911

torch_memory_saver/__init__.py

csrc/torch_memory_saver.cpp

fzyzcjy · 2025-06-15T07:42:35Z

btw I see CUDA graph in your figure, thus #21

fzyzcjy

Good job! Some nits

csrc/torch_memory_saver.cpp

torch_memory_saver/__init__.py

csrc/torch_memory_saver.cpp

fzyzcjy

LGTM, only a bit of nits!

csrc/torch_memory_saver.cpp

fzyzcjy

Only some tiny nits (if you have time maybe spend 3 minute to change and I am ready to merge; if in a hurry I am also ok for current code)

csrc/torch_memory_saver.cpp

torch_memory_saver/__init__.py

Co-authored with: MrAta ([email protected]) ### Checklist Before Starting - [x] Search for similar PR(s). ### What does this PR do? ### Motivation In RL Ecosystem which use colocate design like [verl](https://github.com/volcengine/verl/tree/main), we need to offload training model and load serving model & KV Cache frequently. #### Background - Currently SGLang is using [torch_memory_saver](https://github.com/fzyzcjy/torch_memory_saver) to pause and resume. - [torch_memory_saver](https://github.com/fzyzcjy/torch_memory_saver) is a open source repo that provided easy to use api to hack **cudaMalloc** and **cudaFree** to make sure the virtual address could be consistent after pause and resume, which is critical to ensure CUDA Graph work. - CUDA Graph is critical to make sure SGLang runs faster in decoding phases. #### Here is the current behavior of VERL + SGLang ![Image](https://github.com/user-attachments/assets/e87e7dd6-f223-4de6-8f07-915eb2030ea8) 1. During Training, we have training model and optimizer state in the GPU Memory, and once training is done, we will offload optimizer state to cpu and keep the model weights in GPU, which is needed in Update Weight. 2. During Update Weight, we awake the SGLang engine, so those paused memory of Model Weights and KV Cache will come back. Then we update model from training model to serving model on the fly using the api: `update_weights_in_tensor` 3. After Model being updated, we delete the training model from GPU Memory. Above design works pretty well so far, however, this would waste a big chunk of GPU Memory during rollout, which could cause a few issues we've seen so far: - **Small KV Cache**: We need to use relative lower number of mem fraction ratio (e.g: 0.6), hence our KV Cache has less tokens. Given KV Cache has less tokens, we will hit `RuntimeError: Prefill out of memory. Try to lower your batch size.` when we try prefill large number of requests. - **Out of Memory**: If we use mem fraction ratio 0.8 and run RL for 32B model on 8 H100, it will OOM during update weight #### Challenge - `torch_memory_saver` currently only supports Singleton, hence SGLang will pause and resume KV Cache + Weights together, they are treated as the same group of memory controlled by the singleton `torch_memory_saver` instance #### Proposal ![Image](https://github.com/user-attachments/assets/7fda9638-0dc2-4c14-bc64-cd20616f350f) 1. During Training, we do the same 2. During Update Weight Stage 1, we awake the model weights from SGLang and then update weights 3. During Update Weight Stage 2, we delete the training model weights from GPU Memory 4. Awake the SGLang's KV Cache ![Image](https://github.com/user-attachments/assets/f3dab327-dc2e-4ed8-88d7-15e383f77d25) ### Benefit With above feature, we can train larger model with same GPU, we can also make training/rollout more efficient given we can allocate larger KV Cache ### Solution: Keep using Singleton and provide tag based pause/resume - [x] Support tag based resume/pause: fzyzcjy/torch_memory_saver#20 - [x] Support Multiple Stage Awake in SGLang: sgl-project/sglang#7099 - [ ] Support Multiple Stage Awake in verl: #1911 ### High-Level Design > Demonstrate the high-level design if this PR is complex. ### Specific Changes > List the specific changes. ### API > Demonstrate how the API changes if any. ### Usage Example > Provide usage example(s) for easier usage. ```python # Add code snippet or script demonstrating how to use this ``` ### Test ![Screenshot 2025-06-19 at 12 16 19 PM](https://github.com/user-attachments/assets/a95dd57e-43e1-4f28-8a84-003ec5c043fc) ![Screenshot 2025-06-19 at 12 13 14 PM](https://github.com/user-attachments/assets/f1f4a8a8-1845-4fad-9424-5526d4154dd0) ### Additional Info. - **Issue Number**: Fixes issue # or discussion # if any. - **Training**: [Note which backend this PR will affect: FSDP, Megatron, both, or none] - **Inference**: [Note which backend this PR will affect: vLLM, SGLang, both, or none] ### Checklist Before Submitting - [ ] Read the [Contribute Guide](https://github.com/volcengine/verl?tab=readme-ov-file#contribution-guide). - [ ] Apply [pre-commit checks](https://github.com/volcengine/verl?tab=readme-ov-file#code-linting-and-formatting). - [ ] Add `[BREAKING]` to the PR title if it breaks any API. - [ ] Update the documentation about your changes in the [docs](https://github.com/volcengine/verl/tree/main/docs). - [ ] New CI unit test(s) are added to cover the code path. - [ ] Rely on existing unit tests on CI that covers the code path. --------- Co-authored-by: Chayenne <[email protected]>

Co-authored with: MrAta ([email protected]) ### Checklist Before Starting - [x] Search for similar PR(s). ### What does this PR do? ### Motivation In RL Ecosystem which use colocate design like [verl](https://github.com/volcengine/verl/tree/main), we need to offload training model and load serving model & KV Cache frequently. #### Background - Currently SGLang is using [torch_memory_saver](https://github.com/fzyzcjy/torch_memory_saver) to pause and resume. - [torch_memory_saver](https://github.com/fzyzcjy/torch_memory_saver) is a open source repo that provided easy to use api to hack **cudaMalloc** and **cudaFree** to make sure the virtual address could be consistent after pause and resume, which is critical to ensure CUDA Graph work. - CUDA Graph is critical to make sure SGLang runs faster in decoding phases. #### Here is the current behavior of VERL + SGLang ![Image](https://github.com/user-attachments/assets/e87e7dd6-f223-4de6-8f07-915eb2030ea8) 1. During Training, we have training model and optimizer state in the GPU Memory, and once training is done, we will offload optimizer state to cpu and keep the model weights in GPU, which is needed in Update Weight. 2. During Update Weight, we awake the SGLang engine, so those paused memory of Model Weights and KV Cache will come back. Then we update model from training model to serving model on the fly using the api: `update_weights_in_tensor` 3. After Model being updated, we delete the training model from GPU Memory. Above design works pretty well so far, however, this would waste a big chunk of GPU Memory during rollout, which could cause a few issues we've seen so far: - **Small KV Cache**: We need to use relative lower number of mem fraction ratio (e.g: 0.6), hence our KV Cache has less tokens. Given KV Cache has less tokens, we will hit `RuntimeError: Prefill out of memory. Try to lower your batch size.` when we try prefill large number of requests. - **Out of Memory**: If we use mem fraction ratio 0.8 and run RL for 32B model on 8 H100, it will OOM during update weight #### Challenge - `torch_memory_saver` currently only supports Singleton, hence SGLang will pause and resume KV Cache + Weights together, they are treated as the same group of memory controlled by the singleton `torch_memory_saver` instance #### Proposal ![Image](https://github.com/user-attachments/assets/7fda9638-0dc2-4c14-bc64-cd20616f350f) 1. During Training, we do the same 2. During Update Weight Stage 1, we awake the model weights from SGLang and then update weights 3. During Update Weight Stage 2, we delete the training model weights from GPU Memory 4. Awake the SGLang's KV Cache ![Image](https://github.com/user-attachments/assets/f3dab327-dc2e-4ed8-88d7-15e383f77d25) ### Benefit With above feature, we can train larger model with same GPU, we can also make training/rollout more efficient given we can allocate larger KV Cache ### Solution: Keep using Singleton and provide tag based pause/resume - [x] Support tag based resume/pause: fzyzcjy/torch_memory_saver#20 - [x] Support Multiple Stage Awake in SGLang: sgl-project/sglang#7099 - [ ] Support Multiple Stage Awake in verl: volcengine#1911 ### High-Level Design > Demonstrate the high-level design if this PR is complex. ### Specific Changes > List the specific changes. ### API > Demonstrate how the API changes if any. ### Usage Example > Provide usage example(s) for easier usage. ```python # Add code snippet or script demonstrating how to use this ``` ### Test ![Screenshot 2025-06-19 at 12 16 19 PM](https://github.com/user-attachments/assets/a95dd57e-43e1-4f28-8a84-003ec5c043fc) ![Screenshot 2025-06-19 at 12 13 14 PM](https://github.com/user-attachments/assets/f1f4a8a8-1845-4fad-9424-5526d4154dd0) ### Additional Info. - **Issue Number**: Fixes issue # or discussion # if any. - **Training**: [Note which backend this PR will affect: FSDP, Megatron, both, or none] - **Inference**: [Note which backend this PR will affect: vLLM, SGLang, both, or none] ### Checklist Before Submitting - [ ] Read the [Contribute Guide](https://github.com/volcengine/verl?tab=readme-ov-file#contribution-guide). - [ ] Apply [pre-commit checks](https://github.com/volcengine/verl?tab=readme-ov-file#code-linting-and-formatting). - [ ] Add `[BREAKING]` to the PR title if it breaks any API. - [ ] Update the documentation about your changes in the [docs](https://github.com/volcengine/verl/tree/main/docs). - [ ] New CI unit test(s) are added to cover the code path. - [ ] Rely on existing unit tests on CI that covers the code path. --------- Co-authored-by: Chayenne <[email protected]>

fzyzcjy · 2025-07-08T01:57:49Z

realize that, we may need one separate mem pool per tag in this mode. this is because, suppose:

tag a and allocate 1KB tensor x: indeed torch allocate 2MB memory M, and slice it to satisfy x, so we tag M with a
tag b and allocate 1KB tensor y: indeed torch reuse memory M, so tensor y is indeed with tag a

I will also make a pluggable allocator version which may not have this issue though

hebiao064 · 2025-07-08T03:33:04Z

realize that, we may need one separate mem pool per tag in this mode. this is because, suppose:

tag a and allocate 1KB tensor x: indeed torch allocate 2MB memory M, and slice it to satisfy x, so we tag M with a

tag b and allocate 1KB tensor y: indeed torch reuse memory M, so tensor y is indeed with tag a

I will also make a pluggable allocator version which may not have this issue though

Ah I think I do understand the mechanism you mentioned in your comment (just read this blog few days ago: https://zhuanlan.zhihu.com/p/493646010)

But I wonder how did you realize this problem? and what kind of issues are you foreseeing right now? Will it lead to illegal memory issue we recently hit?

Context:
the illegal memory issue was mitigated by these two PRs

fzyzcjy · 2025-07-08T04:55:31Z

But I wonder how did you realize this problem?

b/c today I am refactoring and adding features to torch_memory_saver, and this comes to my mind.

and what kind of issues are you foreseeing right now?

I expect it to have issues like, tensors are wrongly tagged

hebiao064 · 2025-07-09T02:07:04Z

But I wonder how did you realize this problem?

b/c today I am refactoring and adding features to torch_memory_saver, and this comes to my mind.

and what kind of issues are you foreseeing right now?

I expect it to have issues like, tensors are wrongly tagged

I'm happy to fix the pool arrangement issue, but tbh I am not very clear about the next.

Should I create separate pool when we init those tensors (e.g: kv and weight) and make sure we are malloc memory using the pool passed in?

fzyzcjy · 2025-07-09T02:11:19Z

my personal guess is that change

torch_memory_saver/torch_memory_saver/entrypoint.py

Line 87 in 0fa8e03

with torch.cuda.use_mem_pool(self._mem_pool):

to sth like

# init
self._mem_pools = defaultdict(lambda: torch.cuda.MemPool(...))

# region
mem_pool = self._mem_pools[(tag, enable_cpu_backup)]
... others unchanged

Co-authored with: MrAta ([email protected]) ### Checklist Before Starting - [x] Search for similar PR(s). ### What does this PR do? ### Motivation In RL Ecosystem which use colocate design like [verl](https://github.com/volcengine/verl/tree/main), we need to offload training model and load serving model & KV Cache frequently. #### Background - Currently SGLang is using [torch_memory_saver](https://github.com/fzyzcjy/torch_memory_saver) to pause and resume. - [torch_memory_saver](https://github.com/fzyzcjy/torch_memory_saver) is a open source repo that provided easy to use api to hack **cudaMalloc** and **cudaFree** to make sure the virtual address could be consistent after pause and resume, which is critical to ensure CUDA Graph work. - CUDA Graph is critical to make sure SGLang runs faster in decoding phases. #### Here is the current behavior of VERL + SGLang ![Image](https://github.com/user-attachments/assets/e87e7dd6-f223-4de6-8f07-915eb2030ea8) 1. During Training, we have training model and optimizer state in the GPU Memory, and once training is done, we will offload optimizer state to cpu and keep the model weights in GPU, which is needed in Update Weight. 2. During Update Weight, we awake the SGLang engine, so those paused memory of Model Weights and KV Cache will come back. Then we update model from training model to serving model on the fly using the api: `update_weights_in_tensor` 3. After Model being updated, we delete the training model from GPU Memory. Above design works pretty well so far, however, this would waste a big chunk of GPU Memory during rollout, which could cause a few issues we've seen so far: - **Small KV Cache**: We need to use relative lower number of mem fraction ratio (e.g: 0.6), hence our KV Cache has less tokens. Given KV Cache has less tokens, we will hit `RuntimeError: Prefill out of memory. Try to lower your batch size.` when we try prefill large number of requests. - **Out of Memory**: If we use mem fraction ratio 0.8 and run RL for 32B model on 8 H100, it will OOM during update weight #### Challenge - `torch_memory_saver` currently only supports Singleton, hence SGLang will pause and resume KV Cache + Weights together, they are treated as the same group of memory controlled by the singleton `torch_memory_saver` instance #### Proposal ![Image](https://github.com/user-attachments/assets/7fda9638-0dc2-4c14-bc64-cd20616f350f) 1. During Training, we do the same 2. During Update Weight Stage 1, we awake the model weights from SGLang and then update weights 3. During Update Weight Stage 2, we delete the training model weights from GPU Memory 4. Awake the SGLang's KV Cache ![Image](https://github.com/user-attachments/assets/f3dab327-dc2e-4ed8-88d7-15e383f77d25) ### Benefit With above feature, we can train larger model with same GPU, we can also make training/rollout more efficient given we can allocate larger KV Cache ### Solution: Keep using Singleton and provide tag based pause/resume - [x] Support tag based resume/pause: fzyzcjy/torch_memory_saver#20 - [x] Support Multiple Stage Awake in SGLang: sgl-project/sglang#7099 - [ ] Support Multiple Stage Awake in verl: volcengine#1911 ### High-Level Design > Demonstrate the high-level design if this PR is complex. ### Specific Changes > List the specific changes. ### API > Demonstrate how the API changes if any. ### Usage Example > Provide usage example(s) for easier usage. ```python # Add code snippet or script demonstrating how to use this ``` ### Test ![Screenshot 2025-06-19 at 12 16 19 PM](https://github.com/user-attachments/assets/a95dd57e-43e1-4f28-8a84-003ec5c043fc) ![Screenshot 2025-06-19 at 12 13 14 PM](https://github.com/user-attachments/assets/f1f4a8a8-1845-4fad-9424-5526d4154dd0) ### Additional Info. - **Issue Number**: Fixes issue # or discussion # if any. - **Training**: [Note which backend this PR will affect: FSDP, Megatron, both, or none] - **Inference**: [Note which backend this PR will affect: vLLM, SGLang, both, or none] ### Checklist Before Submitting - [ ] Read the [Contribute Guide](https://github.com/volcengine/verl?tab=readme-ov-file#contribution-guide). - [ ] Apply [pre-commit checks](https://github.com/volcengine/verl?tab=readme-ov-file#code-linting-and-formatting). - [ ] Add `[BREAKING]` to the PR title if it breaks any API. - [ ] Update the documentation about your changes in the [docs](https://github.com/volcengine/verl/tree/main/docs). - [ ] New CI unit test(s) are added to cover the code path. - [ ] Rely on existing unit tests on CI that covers the code path. --------- Co-authored-by: Chayenne <[email protected]>

hebiao064 added 9 commits June 6, 2025 20:31

Support Mulitple Instance

c0d1404

Add CPP change

e331aa0

modify print

4df58ca

fix

fa3d99e

Remove backward compatibility

818f0ba

wrap signature setup into function

9f4eb6a

Fix comments

cea4846

Merge branch 'master' into bhe/support_multiple_instance

06d77ed

Tag based resume

06b1a0a

This was referenced Jun 11, 2025

[RFC] Support Multi-Stage Awake for RL sgl-project/sglang#7009

Closed

Multi-Stage Awake: Support Resume and Pause KV Cache and Weights separately sgl-project/sglang#7099

Merged

hebiao064 marked this pull request as ready for review June 11, 2025 18:35

hebiao064 commented Jun 11, 2025

View reviewed changes

torch_memory_saver/__init__.py Outdated Show resolved Hide resolved

hebiao064 commented Jun 11, 2025

View reviewed changes

torch_memory_saver/__init__.py Outdated Show resolved Hide resolved

hebiao064 commented Jun 11, 2025

View reviewed changes

csrc/torch_memory_saver.cpp Outdated Show resolved Hide resolved

hebiao064 mentioned this pull request Jun 11, 2025

[rollout] feat: Support Multi-stage Awake for SGLang volcengine/verl#1911

Merged

10 tasks

hebiao064 added 2 commits June 12, 2025 01:18

clean up

44c1e83

update readme

873d894

hebiao064 changed the title ~~[Alternative of Multi Instance Solution] Singleton with tag based resume~~ Multi-Stage Awake: Support tag-based Resume and Pause Jun 15, 2025

fzyzcjy mentioned this pull request Jun 15, 2025

Support Mulitple Instance #17

Closed

1 task

fzyzcjy reviewed Jun 15, 2025

View reviewed changes

hebiao064 added 2 commits June 15, 2025 22:30

fix comments

bb2b5f1

Fix comments

5bac911

fzyzcjy reviewed Jun 15, 2025

View reviewed changes

csrc/torch_memory_saver.cpp Outdated Show resolved Hide resolved

fix comment

685460f

fzyzcjy reviewed Jun 16, 2025

View reviewed changes

csrc/torch_memory_saver.cpp Outdated Show resolved Hide resolved

csrc/torch_memory_saver.cpp Outdated Show resolved Hide resolved

csrc/torch_memory_saver.cpp Outdated Show resolved Hide resolved

csrc/torch_memory_saver.cpp Show resolved Hide resolved

fzyzcjy and others added 3 commits June 16, 2025 20:08

Update csrc/torch_memory_saver.cpp

7992adf

Update csrc/torch_memory_saver.cpp

e195f99

fix

2161b06

hebiao064 added 2 commits June 16, 2025 22:14

fix

37ab1f1

cleanup

4fa5c08

fzyzcjy reviewed Jun 17, 2025

View reviewed changes

csrc/torch_memory_saver.cpp Outdated Show resolved Hide resolved

csrc/torch_memory_saver.cpp Outdated Show resolved Hide resolved

csrc/torch_memory_saver.cpp Outdated Show resolved Hide resolved

torch_memory_saver/__init__.py Show resolved Hide resolved

fix

f72a13c

fzyzcjy approved these changes Jun 17, 2025

View reviewed changes

fzyzcjy merged commit f515992 into fzyzcjy:master Jun 17, 2025

hebiao064 mentioned this pull request Jun 18, 2025

Support for multi-instance memory saver #3

Closed

fzyzcjy mentioned this pull request Jul 9, 2025

Question about Hook Modes #32

Open

hebiao064 mentioned this pull request Aug 9, 2025

Support Multi-Stage Awake THUDM/slime#149

Merged

fzyzcjy mentioned this pull request Oct 30, 2025

Fix wrong mode attached to buffers #55

Merged

Multi-Stage Awake: Support tag-based Resume and Pause #20

Multi-Stage Awake: Support tag-based Resume and Pause #20

Uh oh!

Conversation

hebiao064 commented Jun 11, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Motivation

Background

Here is the current behavior of VERL + SGLang

Proposal

Benefit

Execution Plan: Keep using Singleton and provide tag based pause/resume

Uh oh!

Uh oh!

Uh oh!

Uh oh!

fzyzcjy commented Jun 15, 2025

Uh oh!

fzyzcjy left a comment

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

fzyzcjy left a comment

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

fzyzcjy left a comment • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Uh oh!

Uh oh!

Uh oh!

fzyzcjy commented Jul 8, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

hebiao064 commented Jul 8, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

fzyzcjy commented Jul 8, 2025

Uh oh!

hebiao064 commented Jul 9, 2025

Uh oh!

fzyzcjy commented Jul 9, 2025

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

2 participants

hebiao064 commented Jun 11, 2025 •

edited

Loading

fzyzcjy left a comment •

edited

Loading

fzyzcjy commented Jul 8, 2025 •

edited

Loading

hebiao064 commented Jul 8, 2025 •

edited

Loading