Skip to content

[Diffusion] Rollout API: off-loop serialization, spliced msgpack, timing headers, opt-in uint8 video - #36754

Merged
Zhichenzzz merged 5 commits into
sgl-project:sglang-miles-h3from
Rockdu:h3-rollout-perf
Aug 28, 2026
Merged

Zhichenzzz merged 5 commits into
sgl-project:sglang-miles-h3from
Rockdu:h3-rollout-perf

Conversation

@Rockdu

@Rockdu Rockdu commented Aug 28, 2026 •

Copy link
Copy Markdown
Collaborator

What

Five standalone commits on the rollout HTTP path (/rollout/generate):

  1. Off-loop serialization — build + serialize the response in a worker thread (asyncio.to_thread; the safetensors save releases the GIL) and send the pre-built buffers via a fixed-length response with explicit content-length. Identity framing, wire bytes unchanged, zero client change.
  2. msgpack bin-splice — container/bin32 headers are hand-written so large tensor bytes enter the output list by reference instead of being copied through the encoder. b"".join(parts) is byte-identical to msgspec.msgpack.encode(payload) (unit-tested), and the parts are streamed without re-joining, so a serialized tensor is never copied again after safetensors produces it.
  3. x-sgld-timing header — absolute wall-clock marks (srv_recv → msgpack_end) per request, so the trainer can reconstruct a cross-process waterfall. Header only; msgpack body contract untouched.
  4. x-sgld-stages header — the engine's per-stage milliseconds (text encode / denoise / decode), distinguishing a slow denoise from a slow VAE decode client-side.
  5. Opt-in rollout_video_dtype="uint8" — engine-side quantization of the decoded [0,1] video to 0..255 uint8 (consumers divide by 255), ~4x smaller video payload. Default None ships unchanged.

Why

An H3 rollout response carries hundreds of MB of trajectory tensors. Encoding them inline held the GIL and blocked the uvicorn event loop for the full serialization — stalling health checks, reply handling, and next-request dispatch. With the loop free, serialization of request N overlaps the denoise of request N+1 under --sglang-server-concurrency >= 2.

Verified end-to-end on the MiniMax-H3 2-GPU GRPO recipe (3 rollout steps, wire-identical responses, no regression; unit tests in test_msgpack_splice.py / test_rollout_api.py are auto-registered CI).


CI States

Latest PR Test (Base): ❌ Run #33128328860
Latest PR Test (Extra): ❌ Run #33128328737
Latest PR Test (AMD ROCm 7.2): ❌ Run #33128328859

@github-actions github-actions Bot added the diffusion SGLang Diffusion label Aug 28, 2026
@Rockdu
Rockdu marked this pull request as ready for review August 28, 2026 20:40
@Zhichenzzz
Zhichenzzz merged commit aeb02ff into sgl-project:sglang-miles-h3 Aug 28, 2026
80 of 90 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion SGLang Diffusion

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants