docs(cli): correct removed vLLM prefill-worker flag reference - #12581
Merged
Conversation
The KV cache offloading overview described --is-prefill-worker as deprecated. It was removed in v1.4.0 and is absent from the vLLM arg surface in components/src/dynamo/vllm/, so the guidance was accurate in intent but wrong in tense. Follow-up to #12568, which corrected the vLLM configuration reference page but left this occurrence in place. Found while verifying that PR against NVBug 6541824 / DYN-3726. Signed-off-by: Dan Gil <dagil@nvidia.com>
Contributor
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (1)
WalkthroughThe FlexKV disaggregated deployment instructions now identify ChangesFlexKV deployment documentation
Estimated code review effort: 1 (Trivial) | ~2 minutes 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
Comment |
Contributor
This comment has been minimized.
This comment has been minimized.
alec-flowers
approved these changes
Aug 3, 2026
alec-flowers
enabled auto-merge (squash)
August 3, 2026 20:59
Collaborator
Author
|
/ok to test 27b8d75 |
hhzhang16
added a commit
that referenced
this pull request
Aug 4, 2026
dyn-3691-extract-shared-target-pid-cuda-customstorage-operation-layer * 'main' of https://github.com/ai-dynamo/dynamo: (50 commits) docs(cli): correct removed vLLM prefill-worker flag reference (#12581) docs(operator): reserve webhook Ignore for emergencies (#12563) ci(docs): make previews and checks match what actually publishes (#12339) refactor(vllm): organize custom encoder modules (#12416) feat(llm): Select reasoning output field via env var (#11464) feat(runtime): add TLS support to TCP request plane (#10921) fix: convert conditional disagg sglang warning to httperror 400 (#12578) feat(operator): add runtime feature gates (#12421) refactor(runtime): extract PushRouter transport seam behind StreamingDispatch trait (#12447) feat(replay): add deterministic canonical offline reports (#12363) build: bump ModelExpress to 0.5.0(OPS-7978) (#12455) fix(mocker): use logical KV tokens for decode timing (#12583) fix(examples): update Triton example for CUDA 13 + fix libdcgm copy (DYN-3697) (#12577) refactor(operator): implement composition-first DGD reconciliation (#12283) feat(frontend): add basetenkenizer backend (#12376) fix(profiler): configure rapid mocker without planner (#12573) docs(vllm): correct worker-role flags and document --kv-transfer-config (#12568) ci: add Kubernetes deploy test to nightly (#12090) fix(container): reuse pinned protoc in runtime image (#12535) feat(self-host): flip DYN_SELF_HOST_METADATA default to ON (gh-8749) (#11417) ... Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The KV cache offloading overview described
--is-prefill-workeras deprecated. It was removed in v1.4.0 and is absent from the vLLM arg surface incomponents/src/dynamo/vllm/, so the guidance was right in intent but wrong in tense.The mention is kept rather than deleted, because it still helps anyone migrating a script that carries the old flag.
Scope
Follow-up to #12568, which corrected
vllm-configuration.mdxbut left this occurrence. Found while verifying that PR against the originating QA bug.One file, two lines, docs only.
Known gap, deliberately not in this PR
docs/fern/pages/reference/general/releases/deprecations.mdx:342still says these flags "will be removed in a future release," and the ledger's newest section is v1.3.0, so the v1.4.0 removal is unrecorded.That page is mirrored verbatim from the GitHub release notes and its entry counts mirror
RELEASE_STATSinreleases.data.ts. v1.4.0 has not shipped, so writing that section now would invent the source it is meant to mirror. It belongs with the v1.4.0 release notes at GA, tracked separately.Summary by CodeRabbit
--is-prefill-workerflag was removed in v1.4.0.