Skip to content

fix(sglang): reject media for text-only engines - #15520

Merged
furionw merged 8 commits into
ai-dynamo:mainfrom
neilkg:neil/text-only-media-admission-20261001
Oct 7, 2026
Merged

furionw merged 8 commits into
ai-dynamo:mainfrom
neilkg:neil/text-only-media-admission-20261001

Conversation

@neilkg

@neilkg neilkg commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Why

SGLang's own HTTP server rejects media for text-only models, but Dynamo calls SGLang's Engine API directly, which skips that check: SGLang drops the image, video, or audio and the request returns a 200 answer that never saw the media. Rejecting in the worker keeps the decision with the engine that knows its capabilities, works behind any frontend version, and adds no model-card field, so upgrades do not split worker sets by checksum.

Effect

POST /v1/chat/completions  (text-only model on dynamo.sglang, content includes image_url)
before: 200, the answer ignores the image
after:  400 "Model does not accept image input; received unsupported content type 'image_url'."
        streaming: an error event after the initial 200, as with other worker-side rejections

What Change

  • Aggregated, prefill, and decode handlers reject media when is_multimodal is False.
  • The diffusion LM handler rejects media unconditionally; it never forwards media.
  • Multimodal and unknown engines are unchanged; native Generate passthrough is untouched.

Test Plan

  • Unit tests cover text-only, multimodal, and unknown engines; extracted and raw media; empty lists.
  • Not yet exercised end-to-end on a GPU worker.

🤖 Generated with Claude Code

A text-only chat template can silently discard media input. Publish a text-only capability only when the resolved SGLang engine explicitly reports is_multimodal=False, and reject undeclared media before dispatch in the Rust, SGLang and vLLM frontend paths. Missing or malformed declarations preserve existing behavior; multimodal and unknown engines make no declaration.

Co-authored-by: Neil <neil@reflection.ai>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Travis Wilson <travis@reflection.ai>
Signed-off-by: Neil <neil@reflection.ai>
@neilkg
neilkg requested review from a team as code owners October 1, 2026 18:46
@copy-pr-bot

copy-pr-bot Bot commented Oct 1, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@neilkg
neilkg deployed to external_collaborator October 1, 2026 18:46 — with GitHub Actions Active
@neilkg
neilkg deployed to external_collaborator October 1, 2026 18:46 — with GitHub Actions Active
@coderabbitai

coderabbitai Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: ai-dynamo/dynamo/.coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: a6c3683a-8308-44d3-b2aa-66c7ce207bc1

📥 Commits

Reviewing files that changed from the base of the PR and between ec71339 and 2e93d68.

📒 Files selected for processing (9)
  • components/src/dynamo/frontend/sglang_processor.py
  • components/src/dynamo/frontend/tests/test_multimodal_utils.py
  • components/src/dynamo/frontend/tests/test_sglang_processor_unit.py
  • components/src/dynamo/frontend/utils.py
  • components/src/dynamo/frontend/vllm_processor.py
  • components/src/dynamo/sglang/register.py
  • components/src/dynamo/sglang/tests/test_runtime_metadata.py
  • lib/llm/src/local_model/runtime_config.rs
  • lib/llm/src/preprocessor.rs

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 1 remain after this review.


Walkthrough

The change adds runtime input-modality metadata and validation in SGLang and vLLM frontend processors and the Rust preprocessor. When a modality declaration is available, these checks reject media inputs that it does not include.

Changes

Input modality validation

Layer / File(s) Summary
Runtime modality metadata
lib/llm/src/local_model/runtime_config.rs, components/src/dynamo/sglang/register.py, components/src/dynamo/sglang/tests/test_runtime_metadata.py
The runtime configuration defines the input_modalities key. SGLang publishes ["text"] only when its model configuration sets is_multimodal to False.
Frontend request validation
components/src/dynamo/frontend/utils.py, components/src/dynamo/frontend/sglang_processor.py, components/src/dynamo/frontend/vllm_processor.py, components/src/dynamo/frontend/tests/test_multimodal_utils.py, components/src/dynamo/frontend/tests/test_sglang_processor_unit.py
The processors receive modalities parsed from runtime configuration and check user and tool message content before preprocessing. The utility checks image, audio, and video content when a declaration is available.
Preprocessor media validation
lib/llm/src/local_model/runtime_config.rs, lib/llm/src/preprocessor.rs
The preprocessor checks media modalities against runtime configuration and returns InvalidArgument for an undeclared modality. Tests cover rejecting image input for a text-only declaration.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~20 minutes

Merge Risk: ⚪ Minimal · up to 2e93d

The change rejects media sent to models declared text-only and leaves existing behavior unchanged otherwise. No merge-blocking risk was identified.

🚥 Pre-merge checks | ✅ 3 | ❌ 1 | ❓ 1

❌ Failed checks (1 warning, 1 inconclusive)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description explains the reason, behavior, changes, and test plan. However, it omits the required Related Issues reference. Add a valid issue reference under Related Issues, such as "Closes #1234" or "Closes DYN-1234". If no issue exists, create one and reference it.
Docstring Coverage ❓ Inconclusive Docstring coverage is 51.61% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 31 functions across 8 files. (1 skipped: … Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (3 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly summarizes the main change: rejecting media for text-only SGLang engines. It is concise and specific.
Full details: Docstring Coverage

Explanation

Docstring coverage is 51.61% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 31 functions across 8 files. (1 skipped: 1 too large.)

  • Fix all pre-merge checks with AI
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

👋 Hi neilkg! Thank you for contributing to ai-dynamo/dynamo.

Just a reminder: The NVIDIA Test Github Validation CI runs an essential subset of the testing framework to quickly catch errors.Your PR reviewers may elect to test the changes comprehensively before approving your changes.

🚀

@github-actions github-actions Bot added fix external-contribution Pull request is from an external contributor backend::sglang Relates to the sglang backend frontend `python -m dynamo.frontend` and `dynamo-run in=http|text|grpc` labels Oct 1, 2026
Comment thread components/src/dynamo/sglang/tests/test_runtime_metadata.py Outdated
Comment thread lib/llm/src/local_model/runtime_config.rs Outdated
Include canonical input modality declarations in the model-card checksum so
incompatible workers cannot reuse a frontend processor's admission policy.
Preserve legacy checksums for missing or malformed declarations and ignore
ordering and duplicate entries in valid modality sets.

Cover checksum boundaries and discovery admission, and run SGLang's missing
engine-metadata assertions once outside the flag matrix.

Signed-off-by: Neil <neil@reflection.ai>
@neilkg
neilkg deployed to external_collaborator October 1, 2026 21:21 — with GitHub Actions Active
Comment thread lib/llm/src/model_card.rs Outdated
Keep Null for malformed non-array input and a mixed-type array for the
separate malformed-array path. The string input duplicated the Null case.

Signed-off-by: Neil <neil@reflection.ai>
@neilkg
neilkg deployed to external_collaborator October 2, 2026 02:25 — with GitHub Actions Active
Comment thread components/src/dynamo/sglang/tests/test_runtime_metadata.py Outdated
Signed-off-by: Neil <neil@reflection.ai>
@neilkg
neilkg deployed to external_collaborator October 2, 2026 03:01 — with GitHub Actions Active
Comment thread lib/llm/src/local_model/runtime_config.rs Outdated
Signed-off-by: Neil <neil@reflection.ai>
@neilkg
neilkg deployed to external_collaborator October 2, 2026 03:50 — with GitHub Actions Active

@furionw furionw left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for bringing this up, @neilkg .

The idea seems sensible and I'll bring this to my teammate tomorrow

Comment thread components/src/dynamo/sglang/register.py Outdated
@furionw

furionw commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

/ok to test bd8d601

@furionw

furionw commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

chatted with AI and it seems that there is a rollout issue that we need to think of --

Because the declaration participates in mdcsum(), every text-only SGLang worker gets a new checksum on upgrade. Behind an already-upgraded frontend, same-namespace worker replacement (SLURM, Helm, manual) holds newcomers out of the incumbent cohort until it fully drains, then removes the model until the successor builds (lib/llm/src/discovery/controller.rs ~L530-565). Operator-managed DGD rollouts avoid this through hash-suffixed worker namespaces; elsewhere, rolling workers before the frontend avoids it while the serving frontend still ignores the key. A sentence in the PR body would help operators (same pattern as #15179).

@devin-ai-integration

devin-ai-integration Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

✅ Dynamo PR CI passed — run 37542152048 (attempt 2) on 23b0cdd6cf

Gate checks: ✅ backend-status-check · ✅ deploy-status-check · ✅ dynamo-status-check

Posted automatically by Devin for run 37542152048. Updated on every full-CI run of this PR.

@furionw

furionw commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

@neilkg following up on the third item in my earlier comment (rollout note) with the concrete impact, so it can go into the PR description / release notes.

Which models are affected

The worker publishes input_modalities: ["text"] only when SGLang reports model_config.is_multimodal == False (_get_input_modalities in components/src/dynamo/sglang/register.py). Evaluating SGLang 0.5.21's rule (Dynamo's pin) against each checkpoint's config.json, popular models served by dynamo.sglang split as follows:

  • Affected (text-only, declares ["text"]): DeepSeek-R1 / V3.1 / V3.2 / V4-Flash, Qwen3 (including Qwen3-Coder, Qwen3-Next, Qwen3.8), Kimi-K2, GLM-4.7, GLM-5.3, MiniMax-M2 / M2.7, gpt-oss, Llama 3.x, Nemotron 3 Super, Devstral 2, Granite 4.2, Step 3.5 Flash. Also multimodal checkpoints that SGLang serves with multimodality off: Llama 4 or Gemma 3 without --enable-multimodal, or any model started with --language-model-only.
  • Unaffected (multimodal, nothing published): Qwen3-VL, Qwen3.5, Qwen3.6, Kimi-K2.5 / K2.6 / K3, GLM-4.6V, GLM-5.3-Flash, MiniMax-M3, Gemma 4, Mistral Small 4, Nemotron 3 Nano Omni, DeepSeek-V4.1.
  • dynamo.vllm workers publish nothing, so vLLM-served models are unaffected.

Request experience

For affected models, a chat request containing image_url, video_url, or audio_url that used to return 200 with the media silently dropped now returns 400 (Model does not accept image input; received unsupported content type 'image_url'.) on the Rust preprocessor and on the Python SGLang/vLLM chat processors. Text-only requests are unchanged. This matches native SGLang, which has returned 400 for media sent to text-only models since v0.5.17.

Rollout experience

The declaration participates in mdcsum(), so every affected worker's card checksum changes on upgrade. When workers are replaced in place within one Dynamo namespace, behind a frontend that already includes this change:

  1. New workers are held out of the incumbent worker set and receive no traffic (the frontend logs Rejected incompatible workers; the first accepted configuration retains the WorkerSet), while old workers keep serving and capacity shrinks.
  2. When the last old worker leaves, the model is removed and returns 404 until the frontend builds the new worker set.
  3. A frontend restart mid-rollout keeps whichever checksum it discovers first, regardless of how many workers have it.

This is a one-time transition per deployment. Two ways to avoid it:

  • Operator-managed DynamoGraphDeployment (DGD): a DGD rolling update places the new worker generation in a hash-suffixed namespace (DYN_NAMESPACE_WORKER_SUFFIX) while the frontend discovers by prefix (DYN_NAMESPACE_PREFIX), and keeps the previous generation serving until the new one is ready and the old one has drained. No extra action is needed.

  • Without the operator, use a separate namespace for the new workers. Both versions then serve side by side as separate worker sets, and the old ones can be drained afterwards:

    # New-version workers register in namespace "dynamo-v2" (the default namespace is "dynamo").
    DYN_NAMESPACE_WORKER_SUFFIX=v2 python -m dynamo.sglang --model-path zai-org/GLM-5.3 ...
    
    # The frontend must discover both namespaces: omit --namespace (discovers all namespaces),
    # or use a prefix that matches both "dynamo" and "dynamo-v2".
    python -m dynamo.frontend --namespace-prefix dynamo

    Until the old namespace drains, requests are split across the two worker sets in proportion to worker count. For disaggregated deployments, move prefill and decode workers together: a namespace only receives traffic once its worker set is complete.

furionw and others added 2 commits October 6, 2026 15:33
Remove the worker-published input_modalities runtime key, its
model-card checksum entry, and the Rust and Python frontend rejection.
The checksum entry changes the card of every text-only SGLang worker on
upgrade, which splits same-namespace rollouts. The next commit moves the
rejection into the SGLang worker instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: furionw <qiwa@nvidia.com>
SGLang's Engine API skips the HTTP server's text-only media check and
silently drops image, video, and audio input when the model is not
multimodal, so a request returned a 200 answer that never saw its media.

The aggregated, prefill, and decode handlers now raise InvalidArgument
before loading media when the engine reports is_multimodal=False. The
diffusion LM handler never forwards media, so it rejects media
unconditionally. Multimodal and unknown engines are unchanged, and no
model-card field changes, so upgrades do not split worker sets.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: furionw <qiwa@nvidia.com>
@furionw furionw changed the title fix(frontend): reject media for declared text-only models fix(sglang): reject media for text-only engines Oct 6, 2026
@furionw
furionw requested a review from a team as a code owner October 6, 2026 22:38
@furionw
furionw deployed to external_collaborator October 6, 2026 22:38 — with GitHub Actions Active
@furionw

furionw commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

/ok to test b98028e

@furionw

furionw commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

@neilkg heads-up: to land this before today's code freeze, I pushed two commits to your branch that change the approach.

  • f1667dc reverts the input_modalities card declaration, its model-card checksum entry, and the Rust/Python frontend rejection.
  • b98028e moves the rejection into the SGLang worker. The aggregated, prefill, and decode handlers raise InvalidArgument before loading media when the engine reports is_multimodal=False, and the diffusion LM handler, which never forwards media, rejects media unconditionally.

Why: the checksum entry changed the card of every text-only SGLang worker, which splits same-namespace rolling upgrades for most popular text-only models (new workers parked, then a 404 gap). Rejecting in the worker gives the same user-facing fix, a 400 instead of a 200 that silently ignored the media, without a new card contract, and it works behind any frontend version.

Behavior delta vs. your version: non-streaming clients still get the 400. Streaming clients now get an error event after the initial 200 instead of a 400, the same as other worker-side rejections such as TRT-LLM's text-only check. Multimodal and unknown engines are unchanged.

This supersedes my earlier review items and the rollout note above, since there is no checksum change anymore. I also updated the title and description to match. The diagnosis and the original fix are yours. Thank you!

Verification: pre-commit passes, and the new helper logic was exercised locally. The new unit tests (text-only, multimodal, and unknown engines; extracted and raw media; empty lists) run in CI on b98028e. It has not yet been run end-to-end on a GPU worker.

Signed-off-by: furionw <qiwa@nvidia.com>

# Conflicts:
#	components/src/dynamo/sglang/request_handlers/llm/diffusion_handler.py
@furionw
furionw deployed to external_collaborator October 6, 2026 22:41 — with GitHub Actions Active
@furionw

furionw commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

/ok to test 23b0cdd

@furionw

furionw commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

@neilkg one more update: GitHub reported a merge conflict with main, so I merged main into your branch in 23b0cdd. The only conflict was an import line in diffusion_handler.py, where main added thinking_budget_requested next to the new reject_unconsumed_media import; both are kept. No behavior change, and the net diff against main is still the same five SGLang files. Pre-commit passes on the merged files, and CI is authorized for 23b0cdd.

@KrishnanPrash KrishnanPrash left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stamped to unblock

@furionw
furionw enabled auto-merge (squash) October 7, 2026 00:00
@furionw
furionw merged commit 433b4da into ai-dynamo:main Oct 7, 2026
174 of 177 checks passed
@neilkg
neilkg deleted the neil/text-only-media-admission-20261001 branch October 7, 2026 16:00

This branch was successfully deployed

1 active deployment
external_collaborator — 23b0cdd6 Deployed Oct 6, 2026 by furionw via ok-to-test #27452
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backend::sglang Relates to the sglang backend external-contribution Pull request is from an external contributor fix frontend `python -m dynamo.frontend` and `dynamo-run in=http|text|grpc` size/L

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants