Skip to content

docs(cli): correct removed vLLM prefill-worker flag reference - #12581

Merged
alec-flowers merged 3 commits into
mainfrom
dagil-nvidia/docs-dyn3726-flag-removal
Aug 4, 2026
Merged

docs(cli): correct removed vLLM prefill-worker flag reference#12581
alec-flowers merged 3 commits into
mainfrom
dagil-nvidia/docs-dyn3726-flag-removal

Conversation

@dagil-nvidia

@dagil-nvidia dagil-nvidia commented Aug 3, 2026

Copy link
Copy Markdown
Collaborator

Summary

The KV cache offloading overview described --is-prefill-worker as deprecated. It was removed in v1.4.0 and is absent from the vLLM arg surface in components/src/dynamo/vllm/, so the guidance was right in intent but wrong in tense.

-It selects the prefill role with `--disaggregation-mode prefill`; do not use the deprecated
-`--is-prefill-worker` flag.
+It selects the prefill role with `--disaggregation-mode prefill`. The `--is-prefill-worker` flag
+it replaced was removed in v1.4.0.

The mention is kept rather than deleted, because it still helps anyone migrating a script that carries the old flag.

Scope

Follow-up to #12568, which corrected vllm-configuration.mdx but left this occurrence. Found while verifying that PR against the originating QA bug.

One file, two lines, docs only.

Known gap, deliberately not in this PR

docs/fern/pages/reference/general/releases/deprecations.mdx:342 still says these flags "will be removed in a future release," and the ledger's newest section is v1.3.0, so the v1.4.0 removal is unrecorded.

That page is mirrored verbatim from the GitHub release notes and its entry counts mirror RELEASE_STATS in releases.data.ts. v1.4.0 has not shipped, so writing that section now would invent the source it is meant to mirror. It belongs with the v1.4.0 release notes at GA, tracked separately.


Open in Devin Review

Summary by CodeRabbit

  • Documentation
    • Updated FlexKV disaggregated deployment instructions to clarify prefill role selection.
    • Documented that the deprecated --is-prefill-worker flag was removed in v1.4.0.

The KV cache offloading overview described --is-prefill-worker as
deprecated. It was removed in v1.4.0 and is absent from the vLLM arg
surface in components/src/dynamo/vllm/, so the guidance was accurate in
intent but wrong in tense.

Follow-up to #12568, which corrected the vLLM configuration reference
page but left this occurrence in place. Found while verifying that PR
against NVBug 6541824 / DYN-3726.

Signed-off-by: Dan Gil <dagil@nvidia.com>
@dagil-nvidia
dagil-nvidia requested a review from a team as a code owner August 3, 2026 18:31
@github-actions github-actions Bot added docs documentation Improvements or additions to documentation labels Aug 3, 2026

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 1 potential issue.

Open in Devin Review

Comment thread docs/fern/pages/cli/kv-cache-offloading/overview.mdx
@coderabbitai

coderabbitai Bot commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 45ba0a16-fea5-413c-8ba4-6f793b2717aa

📥 Commits

Reviewing files that changed from the base of the PR and between a965876 and 50dab50.

📒 Files selected for processing (1)
  • docs/fern/pages/cli/kv-cache-offloading/overview.mdx

Walkthrough

The FlexKV disaggregated deployment instructions now identify --disaggregation-mode prefill as the prefill role option and state that --is-prefill-worker was removed in v1.4.0.

Changes

FlexKV deployment documentation

Layer / File(s) Summary
Update prefill role instructions
docs/fern/pages/cli/kv-cache-offloading/overview.mdx
The instructions use --disaggregation-mode prefill for the prefill role and state that --is-prefill-worker was removed in v1.4.0.

Estimated code review effort: 1 (Trivial) | ~2 minutes

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly and concisely describes the documentation fix for the removed vLLM prefill-worker flag.
Description check ✅ Passed The description explains the change, scope, rationale, related issue, and deliberate exclusion, so it is mostly complete despite custom headings.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

@datadog-official

This comment has been minimized.

@alec-flowers
alec-flowers enabled auto-merge (squash) August 3, 2026 20:59
@dagil-nvidia

Copy link
Copy Markdown
Collaborator Author

/ok to test 27b8d75

@alec-flowers
alec-flowers merged commit 6463d93 into main Aug 4, 2026
95 checks passed
@alec-flowers
alec-flowers deleted the dagil-nvidia/docs-dyn3726-flag-removal branch August 4, 2026 16:02
hhzhang16 added a commit that referenced this pull request Aug 4, 2026
dyn-3691-extract-shared-target-pid-cuda-customstorage-operation-layer

* 'main' of https://github.com/ai-dynamo/dynamo: (50 commits)
  docs(cli): correct removed vLLM prefill-worker flag reference (#12581)
  docs(operator): reserve webhook Ignore for emergencies (#12563)
  ci(docs): make previews and checks match what actually publishes (#12339)
  refactor(vllm): organize custom encoder modules (#12416)
  feat(llm): Select reasoning output field via env var (#11464)
  feat(runtime): add TLS support to TCP request plane (#10921)
  fix: convert conditional disagg sglang warning to httperror 400 (#12578)
  feat(operator): add runtime feature gates (#12421)
  refactor(runtime): extract PushRouter transport seam behind StreamingDispatch trait (#12447)
  feat(replay): add deterministic canonical offline reports (#12363)
  build: bump ModelExpress to 0.5.0(OPS-7978) (#12455)
  fix(mocker): use logical KV tokens for decode timing (#12583)
  fix(examples): update Triton example for CUDA 13 + fix libdcgm copy (DYN-3697) (#12577)
  refactor(operator): implement composition-first DGD reconciliation (#12283)
  feat(frontend): add basetenkenizer backend (#12376)
  fix(profiler): configure rapid mocker without planner (#12573)
  docs(vllm): correct worker-role flags and document --kv-transfer-config (#12568)
  ci: add Kubernetes deploy test to nightly (#12090)
  fix(container): reuse pinned protoc in runtime image (#12535)
  feat(self-host): flip DYN_SELF_HOST_METADATA default to ON (gh-8749) (#11417)
  ...

Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

docs documentation Improvements or additions to documentation size/XS

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants