Skip to content

[MRV2][PP] Defer sampled-result receives - #53948

Open
chengchengpei wants to merge 9 commits into
vllm-project:mainfrom
chengchengpei:perf/defer-pp-sampled-recv
Open

chengchengpei wants to merge 9 commits into
vllm-project:mainfrom
chengchengpei:perf/defer-pp-sampled-recv

Conversation

@chengchengpei

@chengchengpei chengchengpei commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Related to #53810.

On non-last pipeline-parallel stages, sampled-result receives are currently
posted as soon as their buffers are allocated, although the results are not
consumed until a later scheduler step. For small broadcasts, the receiver-side
NCCL kernel can remain resident while model kernels execute.

This change delays those receives on CUDA Model Runner V2 workers until the
last safe pipeline step and places them after that step's model work. The last
PP rank remains the immediate broadcast source. A result produced at step T
is still consumed at T + PP; only the receiver launch moves.

Activation and scope

The behavior is automatic when all of these are true:

  • Model Runner V2 is active;
  • async scheduling is enabled;
  • pipeline parallel size is greater than one;
  • the platform is CUDA;
  • JIT kernel warmup is enabled and complete; and
  • neither a worker profiler nor multimodal encoder timing is configured.

Synchronous scheduling, CPU, ROCm, single-stage execution, eager execution
without JIT warmup, and diagnostic modes that may synchronize the device retain
immediate receive posting. There are no environment-variable gates.
async_scheduling is resolved from None to a concrete boolean during
VllmConfig validation, before PPHandler is constructed.

Correctness and lifecycle guarantees

  • Warmup retains immediate collective timing; receive deferral starts only
    after warmup finishes.
  • Sampled tokens, sample/rejection counts, and speculative draft tokens retain
    FIFO collective order.
  • The last rank snapshots speculative draft payloads on its main stream before
    queuing them on the potentially backlogged broadcast stream, so later steps
    cannot overwrite an in-flight payload.
  • The sender stays immediate to avoid a cycle with inter-stage activation
    traffic. Its broadcasts are serialized on one dedicated stream; receiver
    deferral can leave the head send waiting, but does not make all queued sends
    simultaneously resident.
  • A consume-time fallback posts any receive that missed its scheduled launch.
  • Pending receives are flushed at PP-group-uniform idle, shutdown, and
    pause/sleep boundaries before device synchronization. The zero-token worker
    flush and model-runner launch are intentionally redundant: the latter also
    protects direct runner invocations.
  • CPU stream placeholders use an explicit launched flag because their event
    recording is a no-op.
  • Rows invalidated while a receive is deferred get both an idx_mapping=-1
    sentinel and a zero accepted-token count, preventing model-specific state
    handlers such as RecoverSSM from committing stale state.
  • PP-only post-update kernels are warmed for every decode shape before serving,
    including Mamba sampled-count/alignment kernels and RecoverSSM commit
    specializations. The all--1 mapping and zero accepted counts make this
    warmup state-neutral.

Sender-side validation envelope

The H100 W1 runs used CUDA 13.3 and NCCL 2.30.7 with TP8/PP4, concurrency 710,
and no speculative draft tokens. The two per-step broadcasts were about 5.7
KiB each. NCCL protocol selection was left automatic and was not printed by
the logs, so this PR does not claim a specific LL/LL128/Simple protocol or a
validated payload-size cutoff. Larger speculative payloads retain the same
ordering and flush invariants in tests, but still need hardware performance
validation.

Validation

Focused tests cover receive cadence, immediate/deferred launch idempotency,
FIFO recovery, speculative draft propagation, invalidated-row masking, idle
flushing, CPU no-op events, warmup specialization and state neutrality, and
pause-before-synchronize ordering.

A matched same-base H100 TP8/PP4 A/B was run on a behavior-equivalent
no-speculation revision of this PR with 4,000 requests, and isolated cold caches. Both arms completed
4,000/4,000 requests with zero failures:

Receive mode Output tok/s Total tok/s Mean TTFT Mean TPOT
Forced immediate after warmup 2,431.4 13,933.9 5,622.0 ms 245.8 ms
Automatic deferred (this PR) 3,368.2 19,302.1 3,303.8 ms 182.5 ms

Automatic deferral improved output and total throughput by 38.5%, reduced
mean TTFT by 41.2%, and reduced mean TPOT by 25.7%. The deferred run had no
errors, tracebacks, NCCL warnings, or CUDA errors. This run also exercised the
immediate last-rank sender under the validated payload envelope without a hang
or observed throughput regression.

The current head additionally hardens mapping dtypes, logical padded-buffer
views, speculative draft snapshots, runtime-sync gating, RecoverSSM warmup, and
invalidated-row handling. Those safeguards are covered by focused tests and
static checks; they have not yet been rerun in the cluster A/B.

AI assistance was used to analyze failures and prepare the changes; the final
implementation and results should be reviewed by the author.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@chengchengpei
chengchengpei marked this pull request as draft August 26, 2026 20:12
@mergify mergify Bot added the mrv2 Model Runner V2 specific label Aug 26, 2026
@chengchengpei
chengchengpei force-pushed the perf/defer-pp-sampled-recv branch from 9cdea89 to 62ff6ff Compare August 26, 2026 21:31
@mergify

mergify Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @chengchengpei.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Sep 1, 2026
@chengchengpei
chengchengpei force-pushed the perf/defer-pp-sampled-recv branch from 929df86 to bccca9f Compare September 5, 2026 04:10
@coderabbitai

coderabbitai Bot commented Sep 5, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Team

Run ID: 355b4657-30a9-4393-900a-dfbd9a6f69cf

📥 Commits

Reviewing files that changed from the base of the PR and between 42c50a73c4a25d4c14572936a010fd24a1608214 and 5e19e157001a5a68c6f91d9af26db15323857357.

📒 Files selected for processing (2)
  • tests/test_config.py
  • vllm/config/vllm.py

Included review availability: Your plan provides up to 10 included reviews per hour; 9 remain after this review.


Walkthrough

Adds configurable deferred pipeline-parallel sampled-token receives. The change validates environment settings, queues and schedules receives, handles idle and shutdown flushing, integrates GPU execution paths, and adds configuration, environment, and PP utility tests.

Changes

Deferred PP sampled-token receives

Layer / File(s) Summary
Configuration and environment contracts
vllm/envs.py, vllm/config/vllm.py, tests/test_envs.py, tests/test_config.py
Adds three environment settings, integer parsing, configuration validation, Qwen4Exp breakable CUDA-graph defaults, and tests for valid and invalid combinations.
Deferred receive queue and synchronization
vllm/v1/worker/gpu/pp_utils.py, tests/v1/worker/test_pp_utils.py
Adds delayed receive queues, post-model launches, idle and explicit flushes, fallback launching, synchronization, statistics, and lifecycle tests.
GPU execution integration
vllm/v1/worker/gpu/model_runner.py, vllm/v1/worker/gpu_worker.py
Enables deferred collectives after warmup, launches receives around model execution, flushes idle and shutdown collectives, and logs statistics.

Estimated code review effort: 4 (Complex) | ~45 minutes

Suggested reviewers: aoshen02, zjy0516

Sequence Diagram(s)

sequenceDiagram
  participant GPUModelRunner
  participant PPHandler
  participant PPBroadcast
  GPUModelRunner->>PPHandler: queue receive slot
  PPHandler->>PPHandler: advance receive queue
  PPHandler->>PPBroadcast: launch sampled, combined, and draft broadcasts
  PPHandler->>GPUModelRunner: wait for receive event
Loading

Merge Risk: ⚪ Minimal · up to f426a

This adds an opt-in deferred pipeline-parallel receive path while retaining immediate receives by default. Supported configurations are validated and the covered lifecycle paths show no current merge-blocking risk.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 29.09% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 55 functions across 8 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly summarizes the main change: deferring sampled-result receives for pipeline parallelism.
Description check ✅ Passed The description explains the receive-deferral behavior, activation conditions, correctness guarantees, and validation results. It is directly related to the changeset.
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Create stacked PR
  • Commit on current branch

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@chengchengpei
chengchengpei force-pushed the perf/defer-pp-sampled-recv branch 2 times, most recently from 5756974 to d0096e6 Compare September 5, 2026 04:30
@mergify mergify Bot removed the needs-rebase label Sep 5, 2026
@chengchengpei
chengchengpei marked this pull request as ready for review September 5, 2026 16:57

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify

mergify Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @chengchengpei.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify

mergify Bot commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @chengchengpei.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Sep 14, 2026
@chengchengpei
chengchengpei force-pushed the perf/defer-pp-sampled-recv branch from 7212070 to 32db8ce Compare September 15, 2026 14:28
@chengchengpei
chengchengpei force-pushed the perf/defer-pp-sampled-recv branch from 2d1b125 to 1839e51 Compare September 29, 2026 19:18
@chengchengpei

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #91940 for commit 1839e51f098d.

@yewentao256 yewentao256 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, Generally LGTM, also want an approval from @njhill or @WoosukKwon

Comment thread vllm/v1/worker/gpu/model_runner.py Outdated
@mergify

mergify Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @chengchengpei.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Oct 1, 2026
Signed-off-by: Chengcheng Pei <5881383+chengchengpei@users.noreply.github.com>
Signed-off-by: Chengcheng Pei <5881383+chengchengpei@users.noreply.github.com>
Signed-off-by: Chengcheng Pei <5881383+chengchengpei@users.noreply.github.com>
Signed-off-by: Chengcheng Pei <5881383+chengchengpei@users.noreply.github.com>
Signed-off-by: Chengcheng Pei <5881383+chengchengpei@users.noreply.github.com>
Signed-off-by: Chengcheng Pei <5881383+chengchengpei@users.noreply.github.com>
Signed-off-by: Chengcheng Pei <5881383+chengchengpei@users.noreply.github.com>
@chengchengpei
chengchengpei force-pushed the perf/defer-pp-sampled-recv branch from 1839e51 to 20d6804 Compare October 3, 2026 08:26
@mergify mergify Bot removed the needs-rebase label Oct 3, 2026
@chengchengpei

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #92748 for commit 20d680484ec7.

@chengchengpei

Copy link
Copy Markdown
Contributor Author

Thanks, Generally LGTM, also want an approval from @njhill or @WoosukKwon

@yewentao256 @njhill @WoosukKwon
All comments addressed. tested.

Can someone approve? thanks

Keep the state-neutral post-update warmup in warmup.py and preserve both upstream prefill and deferred-receive regression coverage after rebasing.

Co-authored-by: Codex <noreply@openai.com>
Signed-off-by: Chengcheng Pei <5881383+chengchengpei@users.noreply.github.com>
@chengchengpei
chengchengpei force-pushed the perf/defer-pp-sampled-recv branch from 20d6804 to c8d9fb1 Compare October 3, 2026 08:53
@vllm-agent

vllm-agent commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

CI selector (shadow): 133 test steps (174 jobs) instead of 75 (91 jobs)

Shadow mode: this changes nothing about what CI runs. It shows what the evidence-based selector would pick for this PR, next to today's rules. How it works.

Feedback welcome: reply here if it would skip a step this change needs, or runs something unrelated.

steps (jobs) Today's rules Selector Would skip Would add
NVIDIA, CPU and others 75 (91) 133 (174) 14 (20) 72 (103)
AMD mirrors 65 (78) 121 (166) 8 (11) 64 (99)
Selector would run (133)
  • amd-kernels-mi355
  • amd-lm-eval-small-models-harness
  • async-engine-inputs-utils-worker
  • basic-correctness ×2
  • basic-correctness-cpu-offload
  • basic-correctness-cumem
  • basic-correctness-prefetch-offload
  • basic-correctness-sleep-mode
  • basic-models-tests-extra-initialization ×14
  • basic-models-tests-initialization
  • basic-models-tests-other
  • batch-invariance-b200
  • batch-invariance-h100
  • benchmarks-cli-test
  • cpu-language-generation-and-pooling-model-tests ×3
  • cpu-multimodal-config
  • cpu-spec-decode-tests
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus
  • cudagraph
  • distributed-comm-ops
  • distributed-compile-comm-4-gpus
  • distributed-compile-rpc-tests-2-gpus
  • distributed-dp-tests-2-gpus
  • distributed-dp-tests-4-gpus
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus
  • distributed-model-tests-2-gpus ×3
  • distributed-mooncakeconnector-pd-accuracy-4-gpus
  • distributed-nixlconnector-pd-accuracy-4-gpus
  • distributed-tests-8xh100
  • distributed-torchrun-examples-4-gpus
  • distributed-torchrun-shutdown-tests-2-gpus
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus
  • e2e-core-1-gpu
  • e2e-core-large-memory
  • e2e-scheduling-1-gpu
  • e2e-scheduling-accuracy-1-gpu
  • elastic-ep-scaling-test
  • engine
  • engine-1-gpu
  • entrypoints-integration-api-server ×4
  • entrypoints-integration-api-server-generate
  • entrypoints-integration-api-server-openai-chat_completion
  • entrypoints-integration-api-server-openai-completion
  • entrypoints-integration-llm
  • entrypoints-integration-multimodal
  • entrypoints-integration-pooling
  • entrypoints-integration-responses-api
  • entrypoints-integration-speech_to_text
  • entrypoints-unit-tests
  • examples
  • extract-hidden-states-integration
  • extract-hidden-states-integration-2-gpus
  • fault-tolerance-e2e-2xh100
  • fusion-e2e-config-sweep-h100
  • fusion-e2e-quick-h100
  • fusion-e2e-tp2-ar-rms-config-sweep-h100
  • fusion-e2e-tp2-asynctp-config-sweep-h100
  • fusion-e2e-tp2-b200
  • fusion-e2e-tp2-quick-h100
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus
  • kernels-b200 ×3
  • kernels-deepgemm-test-h100
  • kernels-root-misc-test-b200
  • kimi-k3-prefix-cache-4xb200 ×2
  • kv-offload-large
  • kv-offload-medium
  • kv-offload-small
  • language-models-tests-extra-standard ×2
  • language-models-tests-granite-l4-compatibility
  • language-models-tests-hybrid ×2
  • language-models-tests-standard
  • lm-eval-dspark-watermark-2xh100
  • lm-eval-small-models
  • lm-eval-turboquant-k3v4nc
  • lm-eval-turboquant-k8v4
  • lm-eval-turboquant-t3nc
  • lm-eval-turboquant-t4nc
  • lm-eval-watermarking
  • lora ×4
  • lora-tp-distributed ×4
  • metrics-tracing-2-gpus
  • model-executor
  • mooncake-ec-tcp-e2e-2-gpus
  • mrcr-eval-small-models
  • multi-modal-accuracy-eval-small-models
  • multi-modal-models-standard-1-qwen2
  • multi-modal-models-standard-2-qwen3-gemma
  • multi-modal-models-standard-3-llava-qwen2-vl
  • multi-modal-models-standard-4-other-whisper
  • multiconnector-nixl-offloading-pd-accuracy-2-gpus
  • multiconnector-nixl-offloading-pd-edge-cases-2-gpus
  • nixlconnector-pd-edge-cases-2-gpus
  • openai-api-correctness
  • pipeline-context-parallelism-4-gpus
  • plugin-tests-2-gpus
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus
  • pytorch-compilation-dynamic-shapes
  • pytorch-compilation-unit-tests
  • pytorch-compilation-unit-tests-h100
  • pytorch-fullgraph-cudagraph-l4-compatibility
  • pytorch-fullgraph-test
  • quantization ×4
  • quantized-models-test
  • quantized-moe-test-b200
  • rayexecutorv2-4-gpus
  • regression
  • replayssm-e2e
  • rust-frontend-core-correctness
  • rust-frontend-distributed
  • rust-frontend-openai-coverage
  • rust-frontend-serve-admin-coverage
  • rust-frontend-tool-use
  • samplers-multimodal-beam-search
  • samplers-test
  • scale-out-ec-e2e-2-gpus
  • sharded-rdt-weight-transfer
  • spec-decode-draft-model ×4
  • spec-decode-eagle-1-deepseek-qwen
  • spec-decode-eagle-2-llama3-qwen-vl-other
  • spec-decode-mtp-deepseek-mimo
  • spec-decode-mtp-gemma4
  • spec-decode-mtp-qwen3-5
  • spec-decode-ngram-suffix
  • spec-decode-speculators
  • v1-core
  • v1-executor-worker
  • v1-kv-connectors ×4
  • v1-logits-oracle
  • v1-metrics-lmeval
  • v1-others-cpu
  • v1-sample
  • v1-spec-decode
Would skip (today's rules run them) (14)
  • ascend-npu-test
  • basic-models-test-other-cpu
  • cpu-distributed-tests-dp-tp
  • cpu-distributed-tests-pp-tp
  • cpu-params-env-tokenizers-parser
  • cpu-reasoning-renderers
  • cpu-tool-parsers
  • jit-monitor-no-runtime-jit
  • kernels-fla-ops-test-b200
  • kernels-mhc-test-b200
  • multi-modal-processor ×4
  • multi-modal-processor-cpu ×4
  • pytorch-compilation-passes-unit-tests
  • v1-kv-offload
Would add (today's rules do not run them) (72)
  • amd-kernels-mi355 (Python record)
  • amd-lm-eval-small-models-harness (Python record)
  • basic-models-tests-extra-initialization ×14 (code map)
  • batch-invariance-b200 (Python record)
  • batch-invariance-h100 (Python record)
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • cudagraph (Python record)
  • distributed-comm-ops (code map)
  • distributed-compile-comm-4-gpus (Python record)
  • distributed-dp-tests-4-gpus (code map)
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus (Python record)
  • distributed-model-tests-2-gpus ×3 (code map)
  • distributed-mooncakeconnector-pd-accuracy-4-gpus (Python record)
  • distributed-nixlconnector-pd-accuracy-4-gpus (Python record)
  • distributed-torchrun-examples-4-gpus (Python record)
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • elastic-ep-scaling-test (Python record)
  • engine (Python record)
  • entrypoints-unit-tests (Python record)
  • examples (Python record)
  • extract-hidden-states-integration (Python record)
  • extract-hidden-states-integration-2-gpus (Python record)
  • fusion-e2e-config-sweep-h100 (Python record)
  • fusion-e2e-quick-h100 (Python record)
  • fusion-e2e-tp2-ar-rms-config-sweep-h100 (Python record)
  • fusion-e2e-tp2-asynctp-config-sweep-h100 (Python record)
  • fusion-e2e-tp2-b200 (Python record)
  • fusion-e2e-tp2-quick-h100 (Python record)
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus (Python record)
  • kernels-b200 ×3 (Python record)
  • kernels-deepgemm-test-h100 (Python record)
  • kimi-k3-prefix-cache-4xb200 ×2 (Python record)
  • kv-offload-large (Python record)
  • kv-offload-medium (Python record)
  • kv-offload-small (Python record)
  • language-models-tests-extra-standard ×2 (code map)
  • lm-eval-small-models (Python record)
  • lm-eval-turboquant-k3v4nc (Python record)
  • lm-eval-turboquant-k8v4 (Python record)
  • lm-eval-turboquant-t3nc (Python record)
  • lm-eval-turboquant-t4nc (Python record)
  • lora ×4 (code map)
  • lora-tp-distributed ×4 (Python record)
  • model-executor (code map)
  • mrcr-eval-small-models (Python record)
  • multi-modal-accuracy-eval-small-models (Python record)
  • multiconnector-nixl-offloading-pd-accuracy-2-gpus (Python record)
  • multiconnector-nixl-offloading-pd-edge-cases-2-gpus (Python record)
  • nixlconnector-pd-edge-cases-2-gpus (Python record)
  • openai-api-correctness (Python record)
  • plugin-tests-2-gpus (Python record)
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus (Python record)
  • quantization ×4 (Python record)
  • quantized-models-test (Python record)
  • quantized-moe-test-b200 (Python record)
  • rayexecutorv2-4-gpus (Python record)
  • rust-frontend-core-correctness (Python record)
  • rust-frontend-openai-coverage (Python record)
  • rust-frontend-tool-use (Python record)
  • samplers-multimodal-beam-search (Python record)
  • samplers-test (Python record)
  • scale-out-ec-e2e-2-gpus (Python record)
  • sharded-rdt-weight-transfer (Python record)
  • spec-decode-draft-model ×4 (Python record)
  • spec-decode-eagle-1-deepseek-qwen (Python record)
  • spec-decode-eagle-2-llama3-qwen-vl-other (Python record)
  • spec-decode-mtp-deepseek-mimo (Python record)
  • spec-decode-mtp-gemma4 (Python record)
  • spec-decode-mtp-qwen3-5 (Python record)
  • spec-decode-ngram-suffix (Python record)
  • spec-decode-speculators (Python record)
AMD mirrors: would skip (8)
  • fault-tolerance-e2e-2xh100
  • kernels-fla-ops-test-b200
  • kernels-mhc-test-b200
  • multi-modal-processor ×4
  • platform-tests
  • pytorch-compilation-passes-unit-tests
  • pytorch-fullgraph-cudagraph-l4-compatibility
  • v1-kv-offload
AMD mirrors: would add (64)
  • basic-models-tests-extra-initialization ×14 (code map)
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • cudagraph (Python record)
  • distributed-comm-ops (code map)
  • distributed-dp-tests-4-gpus (code map)
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus (Python record)
  • distributed-mooncakeconnector-pd-accuracy-4-gpus (Python record)
  • distributed-nixlconnector-pd-accuracy-4-gpus (Python record)
  • distributed-torchrun-examples-4-gpus (Python record)
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • engine (Python record)
  • entrypoints-unit-tests (Python record)
  • examples (Python record)
  • extract-hidden-states-integration (Python record)
  • extract-hidden-states-integration-2-gpus (Python record)
  • fusion-and-compile-unit-tests-2xb200 (Python record)
  • fusion-e2e-config-sweep-h100 (Python record)
  • fusion-e2e-quick-h100 (Python record)
  • fusion-e2e-tp2-ar-rms-config-sweep-h100 (Python record)
  • fusion-e2e-tp2-asynctp-config-sweep-h100 (Python record)
  • fusion-e2e-tp2-b200 (Python record)
  • fusion-e2e-tp2-quick-h100 (Python record)
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus (Python record)
  • kernels-moe-test ×5 (Python record)
  • kernels-quantization-test ×6 (Python record)
  • kv-offload-large (Python record)
  • kv-offload-medium (Python record)
  • kv-offload-small (Python record)
  • language-models-tests-extra-standard ×2 (code map)
  • lm-eval-small-models (Python record)
  • lm-eval-turboquant-k3v4nc (Python record)
  • lm-eval-turboquant-k8v4 (Python record)
  • lm-eval-turboquant-t3nc (Python record)
  • lm-eval-turboquant-t4nc (Python record)
  • lora ×4 (code map)
  • lora-tp-distributed ×4 (Python record)
  • model-executor (code map)
  • mrcr-eval-small-models (Python record)
  • multi-modal-accuracy-eval-small-models (Python record)
  • multiconnector-nixl-offloading-pd-accuracy-2-gpus (Python record)
  • multiconnector-nixl-offloading-pd-edge-cases-2-gpus (Python record)
  • nixlconnector-pd-edge-cases-2-gpus (Python record)
  • nixlconnector-pd-spec-decode-acceptance-2-gpus (Python record)
  • openai-api-correctness (Python record)
  • plugin-tests-2-gpus (Python record)
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus (Python record)
  • quantization ×4 (Python record)
  • quantized-models-test (Python record)
  • rust-frontend-core-correctness (Python record)
  • rust-frontend-openai-coverage (Python record)
  • rust-frontend-tool-use (Python record)
  • samplers-multimodal-beam-search (Python record)
  • samplers-test (Python record)
  • scale-out-ec-e2e-2-gpus (Python record)
  • sharded-rdt-weight-transfer (Python record)
  • spec-decode-draft-model ×4 (Python record)
  • spec-decode-eagle-1-deepseek-qwen (Python record)
  • spec-decode-eagle-2-llama3-qwen-vl-other (Python record)
  • spec-decode-mtp-deepseek-mimo (Python record)
  • spec-decode-mtp-gemma4 (Python record)
  • spec-decode-mtp-qwen3-5 (Python record)
  • spec-decode-ngram-suffix (Python record)
  • spec-decode-speculators (Python record)

11 changed files · base 5f30fc7031 · head 81039f0621 · Python record: build 92706 at 6e517b15c1 · kernel record: table 6e517b1 (build 92706), map 6e517b1 · not counted: 10 build steps, 5 A100 steps the generator no longer emits, 115 optional steps the selector would also run

Keep async scheduling opt-in for callers using the original three arguments. Existing explicit GPU configuration still enables deferred receives only on CUDA after sync-free warmup. Cover legacy calls, synchronous scheduling and non-CUDA fallback.

Co-authored-by: Codex <noreply@openai.com>
Signed-off-by: Chengcheng Pei <5881383+chengchengpei@users.noreply.github.com>
@chengchengpei

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #92751 for commit 81039f062170.

@yewentao256 yewentao256 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the work!

I took a further look, the code diff can be further simplified so that it is easier to land, idealy < 400 LOC.

Please also take a check at runtime collective_rpc, I am not sure if there might be a dead lock.

@chengchengpei

chengchengpei commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor Author

Thanks for the work!

I took a further look, the code diff can be further simplified so that it is easier to land, idealy < 400 LOC.

Please also take a check at runtime collective_rpc, I am not sure if there might be a dead lock.

@yewentao256 @njhill

Thanks for the reviews. Repeated rebases and validation have become too costly for me to continue incrementally. I’m pausing active work until the relevant reviewers can consolidate the remaining blocking feedback and agree on the required validation. I can then decide whether to complete one final revision.

@njhill

njhill commented Oct 6, 2026

Copy link
Copy Markdown
Member

Thanks @chengchengpei that sounds like a good plan. Apologies for the delays, there is a bit of a review bottleneck now especially since it's much quicker to create PRs than review them. I will try to look at this properly asap.

My high-level comment is that (as @yewentao256 said) it still needs quite a bit of simplification, but will try to respond with some more concrete suggestions soon.

We did recently merge this fix #58542 which avoids this problem entirely in the prefill-only case, which may be most common.

@chengchengpei

Copy link
Copy Markdown
Contributor Author

Thanks @chengchengpei that sounds like a good plan. Apologies for the delays, there is a bit of a review bottleneck now especially since it's much quicker to create PRs than review them. I will try to look at this properly asap.

My high-level comment is that (as @yewentao256 said) it still needs quite a bit of simplification, but will try to respond with some more concrete suggestions soon.

We did recently merge this fix #58542 which avoids this problem entirely in the prefill-only case, which may be most common.

@njhill

sure. thanks. take your time.

we deploy this PR (with latest main branch, including #58542) internally because this can improve throughput by about 10-40% on our workload. hope it can be merged so we can upgrade our internal deployment.

@mergify

mergify Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @chengchengpei.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Oct 8, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

mrv2 Model Runner V2 specific needs-rebase ready ONLY add when PR is ready to merge/full CI is needed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants