Skip to content

sync: mirror Slime #2432 Straw replay and queue changes - #458

Closed
aoshen02 wants to merge 22 commits into
vllm-project:mainfrom
aoshen02:codex/slime-2432-stacked
Closed

aoshen02 wants to merge 22 commits into
vllm-project:mainfrom
aoshen02:codex/slime-2432-stacked

Conversation

@aoshen02

Copy link
Copy Markdown
Collaborator

Stacked sync

Depends on #456. This PR mirrors Slime #2432 at bbba9465026f5c16064201d3c84d4552dcf7df57; review the incremental diff against fork PR #6. Do not merge before #456.

Changes

  • Remove obsolete Straw queue limits and require straw-queue>=0.1.2.
  • Preserve archive records and queue namespace during train-only replay, validate metadata, and mirror score-centering validation and NCCL expert-transfer fence.
  • Translate docs and register the two new CPU tests in Buildkite; the upstream SGLang-only binary response test remains engine-specific.

Validation

  • 126 targeted CPU tests; 2 fully async SIGKILL/recovery tests; 86 candidate-image-local tests passed.
  • Black, Ruff, isort, compileall, and diff checks passed.
  • Experimental AMD64 image: aosheninferact/vime@sha256:1ad4d005f6c5f6b0d6e38b16e821fa525021bc5d6deeeea644e3047463453113 (not latest).
  • H200 replay E2E and exact-head full six-suite Buildkite CI in progress.

No automatic merge or production image promotion.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
@read-the-docs-community

read-the-docs-community Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a persistent rollout queue and distributed fully-asynchronous rollout transport using the straw library, allowing prompt tasks, partial rollouts, and training batches to be persisted on shared storage. It also implements Score Centering (SC) to stabilize off-policy reinforcement learning, supporting both top-k and top-p sampling, and integrates R3 routing replay with straw for lazy loading of expert routes. Additionally, support for the supa accelerator backend is added, and the legacy train_async.py script is removed in favor of a unified train.py entrypoint. Feedback on the changes suggests using getattr fallbacks when accessing reference.torch_dtype and self.args.routing_replay_prefetch_microbatches to prevent potential AttributeError runtime crashes.

Comment on lines +285 to +288
kwargs = {
"device": (torch.device("cpu") if isinstance(reference, (TensorRef, DiskTensorRef)) else reference.device),
"dtype": (reference.torch_dtype if isinstance(reference, (TensorRef, DiskTensorRef)) else reference.dtype),
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Accessing reference.torch_dtype directly on reference might raise an AttributeError if reference is an instance of straw.tensor.TensorRef and it does not implement torch_dtype (since dtype is stored as a string representation in straw.tensor.TensorRef). To ensure robustness and prevent potential runtime errors, consider using a fallback to map the string dtype to a torch.dtype object using getattr.

Suggested change
kwargs = {
"device": (torch.device("cpu") if isinstance(reference, (TensorRef, DiskTensorRef)) else reference.device),
"dtype": (reference.torch_dtype if isinstance(reference, (TensorRef, DiskTensorRef)) else reference.dtype),
}
kwargs = {
"device": (torch.device("cpu") if isinstance(reference, (TensorRef, DiskTensorRef)) else reference.device),
"dtype": (
getattr(reference, "torch_dtype", None) or getattr(torch, reference.dtype)
if isinstance(reference, (TensorRef, DiskTensorRef))
else reference.dtype
),
}

Comment on lines +343 to +344
if disk_prefetcher is None:
disk_prefetcher = RoutedExpertsMicrobatchPrefetcher(self.args.routing_replay_prefetch_microbatches)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

If routing_replay_prefetch_microbatches is not defined or configured in self.args, accessing self.args.routing_replay_prefetch_microbatches directly will raise an AttributeError. Consider using getattr with a sensible default value (e.g., 1) to prevent potential runtime crashes.

                if disk_prefetcher is None:
                    disk_prefetcher = RoutedExpertsMicrobatchPrefetcher(
                        getattr(self.args, "routing_replay_prefetch_microbatches", 1)
                    )

Signed-off-by: aoshen02 <aoshen@inferact.ai>
@aoshen02
aoshen02 force-pushed the codex/slime-2432-stacked branch from 92cd9f4 to eaa075d Compare September 30, 2026 11:49
@aoshen02

aoshen02 commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator Author

Validation update (2026-09-30): H200-0 candidate-image E2E passed rollout-only archive export, Straw .straw.json train-only replay, and legacy .pt train-only replay (tests/test_qwen2.5_0.5B_debug_rollout_then_train.py, exit 0). The DCO-only amend moved the head to eaa075d7a069b2c18a3e4b9b3ac943335c14d47a; its Git tree was exactly the same 4f3d144a90b4170d9497ba470e0cf63615ecfdc0 as the image-tested pre-amend commit. DCO passes. Buildkite #1381 then exposed shell quoting in two CPU dependency-install commands; that is fixed by signed-off commit 3f48e09f35740a8f4831aa8a488419fdbc986fab. The authoritative exact-head full-matrix candidate CI is Buildkite #1383, with experimental image aosheninferact/vime@sha256:1ad4d005f6c5f6b0d6e38b16e821fa525021bc5d6deeeea644e3047463453113. Its only image/tree difference is CI YAML quoting; model/runtime code is unchanged. Prior #1378–#1382 were canceled/superseded; #1376 is parent #456, not this head.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
@mergify

mergify Bot commented Sep 30, 2026

Copy link
Copy Markdown

⚠️ The sha of the head commit of this PR conflicts with #441. Mergify cannot evaluate rules on this PR. Once #441 is merged or closed, Mergify will resume processing this PR. ⚠️

@aoshen02

Copy link
Copy Markdown
Collaborator Author

Superseded by #441: this exact head (3f48e09) was fast-forwarded into #441. Buildkite #1383 continues as exact-commit experimental-image full-matrix evidence; the final #441 CI outcome will be recorded there. Closing this duplicate PR.

@aoshen02 aoshen02 closed this Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant