[Doc][Misc] Document batch invariance scheduling limitations - #16232
weijinqian0 merged 5 commits into
Conversation
…onfig Add a Scheduling Limitations section to the batch invariance user guide: chunked prefill, prefix caching, and request preemption (eviction and recomputation) are not supported, and chunked prefill, prefix caching, and the KV cache block size must be configured explicitly. Update the online and offline inference examples with the required --no-enable-chunked-prefill, --no-enable-prefix-caching, and --block-size 128 settings, and apply the same scheduling config to the batch invariance e2e test (test_batch_invariant_tp4.py). Signed-off-by: wangx700 <wangxin700@huawei.com>
…ariance guide State explicitly that --block-size 128 must be used together with --no-enable-chunked-prefill (block_size=128 with enable_chunked_prefill=False for offline inference), instead of listing the settings in parallel. Signed-off-by: wangx700 <wangxin700@huawei.com>
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request improves the documentation and configuration consistency for the batch invariance feature in vLLM. It clarifies that certain scheduling optimizations are incompatible with batch invariance and provides explicit configuration guidance to ensure deterministic behavior. The changes ensure that users and automated tests follow the correct setup requirements. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Doc][Misc] Update batch invariance documentation and tests for scheduling limitationsSuggested PR Summary:
### What this PR does / why we need it?
This pull request updates the batch invariance documentation and tests to reflect scheduling limitations. Specifically, it documents that chunked prefill, prefix caching, and request preemption are not supported with batch invariance, and details how to explicitly disable them while setting the KV cache block size to 128. It also updates the end-to-end consistency test configuration accordingly.
However, the reviewer pointed out that the command-line options `--no-enable-chunked-prefill` and `--no-enable-prefix-caching` do not exist in vLLM. Since these features are disabled by default, they should simply be omitted from the startup command rather than explicitly passed as non-existent flags, and the documentation should be updated to reflect this.
### Does this PR introduce _any_ user-facing change?
Yes, it updates the documentation on how to configure vLLM for batch invariance.
### How was this patch tested?
The patch updates the existing end-to-end test `test_batch_invariant_tp4.py` with the correct configuration parameters.| VLLM_BATCH_INVARIANT=1 vllm serve Qwen/Qwen3-8B \ | ||
| --no-enable-chunked-prefill \ | ||
| --no-enable-prefix-caching \ | ||
| --block-size 128 |
There was a problem hiding this comment.
The command-line options --no-enable-chunked-prefill and --no-enable-prefix-caching do not exist in vLLM. Since both features are disabled by default, you do not need to pass any flags to disable them; simply omitting --enable-chunked-prefill and --enable-prefix-caching is sufficient. Passing unrecognized arguments will cause the vllm serve command to fail.
| VLLM_BATCH_INVARIANT=1 vllm serve Qwen/Qwen3-8B \ | |
| --no-enable-chunked-prefill \ | |
| --no-enable-prefix-caching \ | |
| --block-size 128 | |
| VLLM_BATCH_INVARIANT=1 vllm serve Qwen/Qwen3-8B \\ | |
| --block-size 128 |
There was a problem hiding this comment.
--no-enable-chunked-prefill and --no-enable-prefix-caching was been tested ok.
|
|
||
| Chunked prefill, prefix caching, and request preemption (eviction and recomputation) are not supported with batch invariance. | ||
|
|
||
| These scheduling features are not disabled automatically. You must explicitly disable chunked prefill and prefix caching in your configuration, and pair the chunked prefill disabling with a KV cache block size of 128 — pass `--block-size 128` together with `--no-enable-chunked-prefill` when starting the server, or `block_size=128` together with `enable_chunked_prefill=False` for offline inference — as shown in the examples above. |
There was a problem hiding this comment.
Since --no-enable-chunked-prefill and --no-enable-prefix-caching are not valid vLLM CLI options, the explanation should be updated to reflect that these features are disabled by default and should simply be omitted from the startup command.
| These scheduling features are not disabled automatically. You must explicitly disable chunked prefill and prefix caching in your configuration, and pair the chunked prefill disabling with a KV cache block size of 128 — pass `--block-size 128` together with `--no-enable-chunked-prefill` when starting the server, or `block_size=128` together with `enable_chunked_prefill=False` for offline inference — as shown in the examples above. | |
| These scheduling features are not disabled automatically when batch invariance is enabled. You must ensure chunked prefill and prefix caching are disabled in your configuration (they are disabled by default, so do not pass --enable-chunked-prefill or --enable-prefix-caching when starting the server, or explicitly set enable_chunked_prefill=False and enable_prefix_caching=False for offline inference) and pair this with a KV cache block size of 128 (using --block-size 128 or block_size=128). |
There was a problem hiding this comment.
--no-enable-chunked-prefill and --no-enable-prefix-caching are valid vLLM CLI options
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
|
|
||
| ## Scheduling Limitations | ||
|
|
||
| Chunked prefill, prefix caching, and request preemption (eviction and recomputation) are not supported with batch invariance. |
There was a problem hiding this comment.
it's better to explain how to avoid preemption
There was a problem hiding this comment.
OK,how to avoid preemption is added, and can see more with preemption FAQ
ModelCache.get_or_create only forwards a fixed whitelist of marker-level kwargs to VllmRunner, so enable_chunked_prefill and block_size set at the marker level were silently dropped and the runner fell back to its defaults (enable_chunked_prefill=True, block_size=16). Move both settings into extra_kwargs, which is forwarded verbatim to LLM(), so the batch invariance e2e test actually runs with chunked prefill disabled and a 128-token KV block size. Signed-off-by: wangx700 <wangxin700@huawei.com>
…uide Explain that preemption is triggered when the KV cache runs out and that recomputation breaks batch invariance, and suggest lowering max-num-seqs, keeping max-model-len and per-request max_tokens small, and raising gpu-memory-utilization. Point to the KV cache usage logs and the preemption FAQ for verification. Signed-off-by: wangx700 <wangxin700@huawei.com>
The recomputed prefill of a preempted request includes the tokens generated before the preemption, so attention processes them through the prefill path instead of the original decode path, and the P and D computations cannot be aligned. Tighten the wording as well. Signed-off-by: wangx700 <wangxin700@huawei.com>
…hado/vllm-ascend into main_fix_mrv2_eagle3_mamba * 'main_fix_mrv2_eagle3_mamba' of https://github.com/windshado/vllm-ascend: (42 commits) Update vllm_ascend/worker/v2/model_states/mamba_hybrid.py [Feature][Kimi K3 DSPark] Enable TP for context_proj (vllm-project#16344) [BugFix][SpecDecode] Refresh replicated PCP draft graph cache mappings (vllm-project#16300) [Feature][Model] Integrate Triton KeyPool indexing for GLM-5.3-Flash (vllm-project#16253) [BugFix][Offloader] Re-bind params to NZ static buffers after npu_format_cast (vllm-project#15415) [Feature][Model] Integrate AscendC KDA and causal convolution for GLM-5.3-Flash (vllm-project#16251) [Performance][Communicator] Replace per-layer F.pad with cat of a persistent zero block in MoE prepare (vllm-project#16343) [Feature][Operator] Add DeepSeek V4.1 sparse attention operators (vllm-project#16422) [Doc][Misc] Document batch invariance scheduling limitations (vllm-project#16232) [CI][MRV2] Enable mrv2 dspark e2e test (vllm-project#16319) [BugFix] Precast MoE gate weight_fp32 to avoid aclop Cast (vllm-project#16189) [Feature][MRV2][310P] MRv2 adapting MTP on the 310P for Qwen3.5 (vllm-project#16043) [Revert] Revert "[Feature][MRV1][MRV2] Refactor Host-Side Parameter Updates for ACL Graph Replay." (vllm-project#15908) (vllm-project#16409) [Feature][Ops] Add Triton KeyPool compression and pooled indexing (vllm-project#16243) [Feature][Attention] Support NoPE in the shared SFA backend (vllm-project#16252) [Performance][Model] Reuse fused mHC operators for GLM-5.3-Flash (vllm-project#16321) [Feature][Model] Enable MiniMax-M3 FP8 MSA index score on A5 (vllm-project#15918) [Performance][KDA] Reduce preprocessing copies and redundant output masks (vllm-project#16067) [Feature][Model][MTP] Support speculative decoding for GLM-5.3-Flash (vllm-project#16214) [BugFix][Model] Skip unused hash-router bias when loading DeepSeek-V4 weights (vllm-project#16259) ...
…oject#16232) ### What this PR does / why we need it? Documents the scheduling limitations of batch invariance in the user guide: chunked prefill, prefix caching, and request preemption (eviction and recomputation) are not supported. These features are not disabled automatically, so the guide now instructs users to explicitly disable chunked prefill and prefix caching, and pairs `--block-size 128` with the chunked prefill disabling. The online/offline inference examples and the batch invariance e2e test are updated with the same config. ### Does this PR introduce _any_ user-facing change? Documentation-only update plus an e2e test configuration adjustment. ### How was this patch tested? Doc-only change; the e2e test config (`test_batch_invariant_tp4.py`) follows the documented settings. CI verification is sufficient. - VLLM_WORKER_MULTIPROC_METHOD=spawn pytest -sv tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py ```bash [DEBUG] Final model_config: {'model_name': ..Qwen3-30B-A3B', 'max_num_seqs': 144, 'gpu_memory_utilization': 0.95, 'max_model_len': 8192, 'dtype': 'bfloat16', 'tensor_parallel_size': 4, 'enable_prefix_caching': False, 'distributed_executor_backend': 'mp', 'compilation_config': {'cudagraph_capture_sizes': [1, 32, 64]}, 'extra_kwargs': {'load_format': 'dummy', 'hf_overrides': {'num_hidden_layers': 2}, 'enable_chunked_prefill': False, 'block_size': 128}} [DEBUG] vllm_runner fixture - model_config: {'model_name': '..Qwen3-30B-A3B', 'max_num_seqs': 144, 'gpu_memory_utilization': 0.95, 'max_model_len': 8192, 'dtype': 'bfloat16', 'tensor_parallel_size': 4, 'enable_prefix_caching': False, 'distributed_executor_backend': 'mp', 'compilation_config': {'cudagraph_capture_sizes': [1, 32, 64]}, 'extra_kwargs': {'load_format': 'dummy', 'hf_overrides': {'num_hidden_layers': 2}, 'enable_chunked_prefill': False, 'block_size': 128}} ... PASSEDINFO 09-10 20:42:39 [utils.py:620] [shutdown] Process manager: send sigterm to process EngineCore (EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1350] [shutdown] EngineCore: trigger received signal=SIGTERM (EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1501] [shutdown] EngineCore: start mode=abort timeout=0s (EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1532] [shutdown] EngineCore: request processing complete; starting resource teardown (EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1363] [shutdown] EngineCore: exiting busy loop (EngineCore pid=64849) INFO 09-10 20:42:39 [multiproc_executor.py:472] [shutdown] Executor: waiting for worker exit count=4 (Worker_TP0 pid=65187) INFO 09-10 20:42:39 [multiproc_executor.py:836] Parent process exited, terminating worker queues WARNING 09-10 20:42:44 [utils.py:640] [shutdown] Process manager: force killing remaining processes count=1 WARNING 09-10 20:42:44 [utils.py:645] [shutdown] Process manager: force killing remaining process EngineCore pid 64849 (EngineCore pid=64849) WARNING 09-10 20:42:44 [multiproc_executor.py:484] [shutdown] Executor: workers still running after grace period; sending SIGTERM count=3 [INFO] Model cache cleared =============================== warnings summary =============================== ../../../../../usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: 14 warnings /usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`. warnings.warn( tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111 /mnt/share/w00899129/vllm_code_0908/vllm-ascend/tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111: PytestUnknownMarkWarning: Unknown pytest.mark.model - is this a typo? You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html @pytest.mark.model( -- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html ================== 1 passed, 15 warnings in 118.88s (0:01:58) ================== /usr/local/python3.11.10/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' ``` - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: wangx700 <wangxin700@huawei.com> Signed-off-by: tianming2009 <13246728590@163.com>
…oject#16232) ### What this PR does / why we need it? Documents the scheduling limitations of batch invariance in the user guide: chunked prefill, prefix caching, and request preemption (eviction and recomputation) are not supported. These features are not disabled automatically, so the guide now instructs users to explicitly disable chunked prefill and prefix caching, and pairs `--block-size 128` with the chunked prefill disabling. The online/offline inference examples and the batch invariance e2e test are updated with the same config. ### Does this PR introduce _any_ user-facing change? Documentation-only update plus an e2e test configuration adjustment. ### How was this patch tested? Doc-only change; the e2e test config (`test_batch_invariant_tp4.py`) follows the documented settings. CI verification is sufficient. - VLLM_WORKER_MULTIPROC_METHOD=spawn pytest -sv tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py ```bash [DEBUG] Final model_config: {'model_name': ..Qwen3-30B-A3B', 'max_num_seqs': 144, 'gpu_memory_utilization': 0.95, 'max_model_len': 8192, 'dtype': 'bfloat16', 'tensor_parallel_size': 4, 'enable_prefix_caching': False, 'distributed_executor_backend': 'mp', 'compilation_config': {'cudagraph_capture_sizes': [1, 32, 64]}, 'extra_kwargs': {'load_format': 'dummy', 'hf_overrides': {'num_hidden_layers': 2}, 'enable_chunked_prefill': False, 'block_size': 128}} [DEBUG] vllm_runner fixture - model_config: {'model_name': '..Qwen3-30B-A3B', 'max_num_seqs': 144, 'gpu_memory_utilization': 0.95, 'max_model_len': 8192, 'dtype': 'bfloat16', 'tensor_parallel_size': 4, 'enable_prefix_caching': False, 'distributed_executor_backend': 'mp', 'compilation_config': {'cudagraph_capture_sizes': [1, 32, 64]}, 'extra_kwargs': {'load_format': 'dummy', 'hf_overrides': {'num_hidden_layers': 2}, 'enable_chunked_prefill': False, 'block_size': 128}} ... PASSEDINFO 09-10 20:42:39 [utils.py:620] [shutdown] Process manager: send sigterm to process EngineCore (EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1350] [shutdown] EngineCore: trigger received signal=SIGTERM (EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1501] [shutdown] EngineCore: start mode=abort timeout=0s (EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1532] [shutdown] EngineCore: request processing complete; starting resource teardown (EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1363] [shutdown] EngineCore: exiting busy loop (EngineCore pid=64849) INFO 09-10 20:42:39 [multiproc_executor.py:472] [shutdown] Executor: waiting for worker exit count=4 (Worker_TP0 pid=65187) INFO 09-10 20:42:39 [multiproc_executor.py:836] Parent process exited, terminating worker queues WARNING 09-10 20:42:44 [utils.py:640] [shutdown] Process manager: force killing remaining processes count=1 WARNING 09-10 20:42:44 [utils.py:645] [shutdown] Process manager: force killing remaining process EngineCore pid 64849 (EngineCore pid=64849) WARNING 09-10 20:42:44 [multiproc_executor.py:484] [shutdown] Executor: workers still running after grace period; sending SIGTERM count=3 [INFO] Model cache cleared =============================== warnings summary =============================== ../../../../../usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: 14 warnings /usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`. warnings.warn( tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111 /mnt/share/w00899129/vllm_code_0908/vllm-ascend/tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111: PytestUnknownMarkWarning: Unknown pytest.mark.model - is this a typo? You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html @pytest.mark.model( -- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html ================== 1 passed, 15 warnings in 118.88s (0:01:58) ================== /usr/local/python3.11.10/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' ``` - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: wangx700 <wangxin700@huawei.com> Signed-off-by: like-0517 <ithwlike@126.com>
What this PR does / why we need it?
Documents the scheduling limitations of batch invariance in the user guide: chunked prefill, prefix caching, and request preemption (eviction and recomputation) are not supported. These features are not disabled automatically, so the guide now instructs users to explicitly disable chunked prefill and prefix caching, and pairs
--block-size 128with the chunked prefill disabling. The online/offline inference examples and the batch invariance e2e test are updated with the same config.Does this PR introduce any user-facing change?
Documentation-only update plus an e2e test configuration adjustment.
How was this patch tested?
Doc-only change; the e2e test config (
test_batch_invariant_tp4.py) follows the documented settings. CI verification is sufficient.[DEBUG] Final model_config: {'model_name': ..Qwen3-30B-A3B', 'max_num_seqs': 144, 'gpu_memory_utilization': 0.95, 'max_model_len': 8192, 'dtype': 'bfloat16', 'tensor_parallel_size': 4, 'enable_prefix_caching': False, 'distributed_executor_backend': 'mp', 'compilation_config': {'cudagraph_capture_sizes': [1, 32, 64]}, 'extra_kwargs': {'load_format': 'dummy', 'hf_overrides': {'num_hidden_layers': 2}, 'enable_chunked_prefill': False, 'block_size': 128}} [DEBUG] vllm_runner fixture - model_config: {'model_name': '..Qwen3-30B-A3B', 'max_num_seqs': 144, 'gpu_memory_utilization': 0.95, 'max_model_len': 8192, 'dtype': 'bfloat16', 'tensor_parallel_size': 4, 'enable_prefix_caching': False, 'distributed_executor_backend': 'mp', 'compilation_config': {'cudagraph_capture_sizes': [1, 32, 64]}, 'extra_kwargs': {'load_format': 'dummy', 'hf_overrides': {'num_hidden_layers': 2}, 'enable_chunked_prefill': False, 'block_size': 128}} ... PASSEDINFO 09-10 20:42:39 [utils.py:620] [shutdown] Process manager: send sigterm to process EngineCore (EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1350] [shutdown] EngineCore: trigger received signal=SIGTERM (EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1501] [shutdown] EngineCore: start mode=abort timeout=0s (EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1532] [shutdown] EngineCore: request processing complete; starting resource teardown (EngineCore pid=64849) INFO 09-10 20:42:39 [core.py:1363] [shutdown] EngineCore: exiting busy loop (EngineCore pid=64849) INFO 09-10 20:42:39 [multiproc_executor.py:472] [shutdown] Executor: waiting for worker exit count=4 (Worker_TP0 pid=65187) INFO 09-10 20:42:39 [multiproc_executor.py:836] Parent process exited, terminating worker queues WARNING 09-10 20:42:44 [utils.py:640] [shutdown] Process manager: force killing remaining processes count=1 WARNING 09-10 20:42:44 [utils.py:645] [shutdown] Process manager: force killing remaining process EngineCore pid 64849 (EngineCore pid=64849) WARNING 09-10 20:42:44 [multiproc_executor.py:484] [shutdown] Executor: workers still running after grace period; sending SIGTERM count=3 [INFO] Model cache cleared =============================== warnings summary =============================== ../../../../../usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: 14 warnings /usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`. warnings.warn( tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111 /mnt/share/w00899129/vllm_code_0908/vllm-ascend/tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111: PytestUnknownMarkWarning: Unknown pytest.mark.model - is this a typo? You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html @pytest.mark.model( -- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html ================== 1 passed, 15 warnings in 118.88s (0:01:58) ================== /usr/local/python3.11.10/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d '