Skip to content

[Sync] Mirror Slime through 8c17b676 (#2437/#2434) - #441

Open
aoshen02 wants to merge 35 commits into
vllm-project:mainfrom
aoshen02:codex/slime-2393-sync
Open

aoshen02 wants to merge 35 commits into
vllm-project:mainfrom
aoshen02:codex/slime-2393-sync

Conversation

@aoshen02

@aoshen02 aoshen02 commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Mirror Slime from 4c193f1f through 8c17b676: 14 applicable merged PRs;
Ascend #2424 remains explicitly excluded. Biren/SUPA is mirrored as requested.

Includes score centering, Straw-backed distributed fully-async rollout and
checkpoint continuation, PipelineRL flush, and reference full-vocabulary top-p
handling. Existing Vime-native overlays are retained separately from translations.

Engine patches

The vLLM source pin remains e9757321527ca1ecd514c07c1418dd2c53da3d19.

Patch Purpose / PR
vllm-pull_weights.patch Pull-checkpoint transport: aoshen02/vllm#45
vllm.patch Weight-version propagation: vllm-project/vllm#53199; terminal speculative metrics preservation: vllm-project/vllm#55708; empty terminal token-stream events: merged vllm-project/vllm#47933
vllm-score-centering.patch Sampling-mask transport and exact top-p extension: aoshen02/vllm#83
vllm-pd-request-metrics.patch PD request metrics: aoshen02/vllm#29
vllm-inflight-queue-diagnostics.patch Queue diagnostics: vllm-project/vllm#55274
vllm-aux-output-reset.patch Backport merged vllm-project/vllm#59060, absent from the pin

Final audit

  • Reviewed all 150 changed paths against the Slime tip and the pre-sync Vime baseline.
  • Fixed coalesced mask IDs/logprobs and terminal speculative metrics; real-engine regressions fail before and pass after.
  • Validate actual processed_logprobs server configuration, including external engines, rather than a nonexistent environment variable.
  • Removed obsolete per-file R3 spill and its flags/tests; use upstream Straw references. Restore direct Straw documentation translations and omitted upstream comments.
  • Audit fixes net-remove approximately 500 lines. Historical engine gaps outside this sync are not claimed closed.

Validation

  • Relevant post-cleanup CPU regressions: 292 passed; baked runtime image checks: 175 passed.
  • Final image smoke: 3 output-coalescing regressions + 4 documentation checks passed.
  • All six patches apply sequentially to the exact pin; 2,533 patched Python files parse successfully.
  • Changed-file pre-commit passes.
  • Final commit: f32a36f1d8f811260f92c1901e808c35ad13c4ea.
  • Experimental AMD64 image: aosheninferact/vime@sha256:d2d60d7ac5a4fc27299bee5a6ad9f24ba30508f52590dd7f10d4b66516f56cf1.
  • #1445 passed 49 jobs; PipelineRL NCCL with periodic cache flush failed with Straw file-descriptor exhaustion. Exact-image sequential read checks show no FD growth; a deliberately exhausted soft limit reproduces the exception. Two CI startup lines now raise the soft FD limit to the available hard limit and log both; concurrent resource leaks are not ruled out by this check.
  • H200 four-GPU control with soft FD limit1024 reproduces the exact Straw failure; raising the limit allows the normal socket concurrency (observed generation-actor FD count1614). A second high-limit run exposed an independent empty-terminal SSE omission already fixed by upstream #47933; backport that condition rather than weaken the abort assertion. Corrected actual-serving regression matrix: six failures/two passes before, eight passes after; full transport file28 passes. Final image passed two independent complete four-GPU PipelineRL runs (one trainer, three rollout engines, NCCL, NVLS0, flush interval2), including all three training rollouts, weight-version continuity, periodic abort and checkpoint assertions; both exit0 in4m03s.
  • Full final CI #1454: pending; all six GPU suites must pass before merge readiness.
  • Previous #1429 passed 50 jobs on the earlier revision, not evidence of final-head green.

vllm/vime:latest remains unchanged; ARM publication follows full AMD64 CI.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
@read-the-docs-community

read-the-docs-community Bot commented Sep 19, 2026 •

Copy link
Copy Markdown

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request consolidates asynchronous training by removing train_async.py and integrating fully-asynchronous rollout capabilities directly into train.py. It also introduces support for the SUPA accelerator (Biren GPUs), implements disk-spilling and prefetching for routed experts (R3) to optimize host memory usage, and adds a utility to reset the CUDA stack size after model offload. The review feedback highlights several critical robustness improvements, including safely accessing torch.version attributes to prevent AttributeError on standard PyTorch installations, defensively handling potentially None values for args.num_experts to avoid TypeError, and correcting test mocks to avoid attribute errors during unit testing.

Comment on lines +27 to +42
def reset_cuda_stack_size() -> None:
"""Release an enlarged CUDA per-thread stack after model offload."""
if torch.version.cuda is None or torch.version.hip is not None or not torch.cuda.is_initialized():
return
torch.cuda.synchronize()
driver = _cuda_stack_api()
previous = ctypes.c_size_t()
error = driver.cuCtxGetLimit(ctypes.byref(previous), 0) # CU_LIMIT_STACK_SIZE
if error:
raise RuntimeError(f"cuCtxGetLimit(CU_LIMIT_STACK_SIZE) failed: CUDA error {error}")
if previous.value <= 1024:
return
error = driver.cuCtxSetLimit(0, 1024)
if error:
raise RuntimeError(f"cuCtxSetLimit(CU_LIMIT_STACK_SIZE) failed: CUDA error {error}")
logger.info("Reset CUDA stack limit after offload: %d -> 1024 bytes", previous.value)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Accessing torch.version.hip directly will raise an AttributeError on standard (non-ROCm) PyTorch installations because the hip attribute is not defined on the torch.version module in standard builds. To prevent this, use getattr(torch.version, "hip", None) and getattr(torch.version, "cuda", None) to safely check the versions. Additionally, wrapping the entire body of reset_cuda_stack_size in a try...except block is highly recommended to ensure that any unexpected CUDA stack API loading or driver failures do not crash the training process.

def reset_cuda_stack_size() -> None:
    """Release an enlarged CUDA per-thread stack after model offload."""
    cuda_version = getattr(torch.version, "cuda", None)
    hip_version = getattr(torch.version, "hip", None)
    if cuda_version is None or hip_version is not None or not torch.cuda.is_initialized():
        return
    try:
        torch.cuda.synchronize()
        driver = _cuda_stack_api()
        previous = ctypes.c_size_t()
        error = driver.cuCtxGetLimit(ctypes.byref(previous), 0)  # CU_LIMIT_STACK_SIZE
        if error:
            raise RuntimeError(f"cuCtxGetLimit(CU_LIMIT_STACK_SIZE) failed: CUDA error {error}")
        if previous.value <= 1024:
            return
        error = driver.cuCtxSetLimit(0, 1024)
        if error:
            raise RuntimeError(f"cuCtxSetLimit(CU_LIMIT_STACK_SIZE) failed: CUDA error {error}")
        logger.info("Reset CUDA stack limit after offload: %d -> 1024 bytes", previous.value)
    except Exception as exc:
        logger.warning("Failed to reset CUDA stack size: %s", exc)

Comment on lines +14 to +15
monkeypatch.setattr(memory_utils.torch.version, "cuda", cuda)
monkeypatch.setattr(memory_utils.torch.version, "hip", hip)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Using monkeypatch.setattr to set cuda and hip directly on memory_utils.torch.version will raise an AttributeError on standard PyTorch installations because the hip attribute does not exist on the torch.version module. To fix this, mock the entire version attribute as a SimpleNamespace containing both cuda and hip attributes, which is consistent with the other test in this file.

Suggested change
monkeypatch.setattr(memory_utils.torch.version, "cuda", cuda)
monkeypatch.setattr(memory_utils.torch.version, "hip", hip)
monkeypatch.setattr(memory_utils.torch, "version", SimpleNamespace(cuda=cuda, hip=hip))

if moe_layers:
# Cast first so uint8 expert ids compare correctly with num_experts=256.
moe_routes = experts[:, moe_layers, :].to(torch.int64)
num_experts = int(getattr(args, "num_experts", torch.iinfo(torch.int32).max))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

If args.num_experts is explicitly set to None, calling int(None) will raise a TypeError. It is safer to retrieve the value defensively and default to torch.iinfo(torch.int32).max if it is None or not present.

Suggested change
num_experts = int(getattr(args, "num_experts", torch.iinfo(torch.int32).max))
num_experts_val = getattr(args, "num_experts", None)
num_experts = int(num_experts_val) if num_experts_val is not None else torch.iinfo(torch.int32).max

Comment thread vime/utils/routed_experts.py Outdated
if sample.rollout_routed_experts is None:
if sample.loss_mask is None or any(sample.loss_mask):
return sample
dtype = torch.uint8 if args.num_experts <= 256 else torch.int32

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

If args.num_experts is None or not defined on args, accessing args.num_experts directly will raise a TypeError or AttributeError. It is safer to use getattr defensively to check if num_experts is set and is less than or equal to 256.

Suggested change
dtype = torch.uint8 if args.num_experts <= 256 else torch.int32
num_experts = getattr(args, "num_experts", None)
dtype = torch.uint8 if num_experts is not None and num_experts <= 256 else torch.int32

Signed-off-by: aoshen02 <aoshen@inferact.ai>
@aoshen02
aoshen02 marked this pull request as ready for review September 19, 2026 01:28
Signed-off-by: aoshen02 <aoshen@inferact.ai>
@aoshen02
aoshen02 force-pushed the codex/slime-2393-sync branch from 63502a1 to 3f5f82d Compare September 21, 2026 00:46
@aoshen02 aoshen02 changed the title [Sync] Update from Slime through #2393 [Sync] Update from Slime through #2394 Sep 21, 2026
@aoshen02 aoshen02 changed the title [Sync] Update from Slime through #2394 [Sync] Update from Slime through #2407 Sep 24, 2026
@aoshen02
aoshen02 force-pushed the codex/slime-2393-sync branch from a712b9e to b600940 Compare September 24, 2026 01:43
Signed-off-by: aoshen02 <aoshen@inferact.ai>
@aoshen02
aoshen02 force-pushed the codex/slime-2393-sync branch from b600940 to cd3992f Compare September 24, 2026 02:08
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
@aoshen02
aoshen02 force-pushed the codex/slime-2393-sync branch from 17390d4 to 72bdd14 Compare September 28, 2026 04:20
Signed-off-by: aoshen02 <aoshen@inferact.ai>
@aoshen02
aoshen02 force-pushed the codex/slime-2393-sync branch from 3739f33 to 30ce1fe Compare September 28, 2026 06:44
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
@aoshen02
aoshen02 force-pushed the codex/slime-2393-sync branch from 2e2a7f4 to 927a686 Compare September 29, 2026 13:42
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
@aoshen02 aoshen02 changed the title [Sync] Update from Slime through #2407 [Sync] Update from Slime through #2432 Sep 30, 2026
@aoshen02

Copy link
Copy Markdown
Collaborator Author

Exact-head full CI is green: Buildkite #1385 passed all automatic CPU checks and 40/40 GPU jobs on commit 3f48e09f35740a8f4831aa8a488419fdbc986fab using experimental AMD64 image aosheninferact/vime@sha256:1ad4d005f6c5f6b0d6e38b16e821fa525021bc5d6deeeea644e3047463453113. The sole first-attempt failure was Buildkite agent loss (exit_status=-1, no test exception) on Qwen3-30B-A3B R3; one unchanged retry passed. GitHub commit status now points to #1385 and is green. Per the sync SOP, I then checked ARM64 access using ssh -o BatchMode=yes -o ConnectTimeout=5 gb200-rack1-05 'uname -m; hostname'; it timed out during SSH banner exchange. No ARM64 candidate image was built or validated, and vllm/vime:latest remains unchanged.

@aoshen02

Copy link
Copy Markdown
Collaborator Author

ARM64 follow-up: ssh -G gb200-rack1-05 reveals ProxyJump gcp-gb200-head. This alias is not an independent non-GCP route; the user has disallowed use of the GCP GB200 host. I will not build through that route. ARM64 candidate build/validation remains pending an approved accessible ARM64 host.

Port #2437 flush scheduling and #2434 reference log-prob replay fix. Replace request-bound TITO weight version with per-output propagation through vLLM #53199, and preserve exact top-p score-centering support.

Signed-off-by: aoshen02 <aoshen@inferact.ai>
@aoshen02
aoshen02 force-pushed the codex/slime-2393-sync branch from a3e63cf to 34b247c Compare October 4, 2026 01:22
@aoshen02 aoshen02 changed the title [Sync] Update from Slime through #2432 [Sync] Mirror Slime through 8c17b676 (#2437/#2434) Oct 4, 2026
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
@aoshen02
aoshen02 force-pushed the codex/slime-2393-sync branch from a476dd9 to d978296 Compare October 6, 2026 02:15
…bilities

Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
@aoshen02
aoshen02 force-pushed the codex/slime-2393-sync branch from 48d31ee to c5100bd Compare October 6, 2026 03:08
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Signed-off-by: aoshen02 <aoshen@inferact.ai>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant