Skip to content

fix(frontend): propagate worker InvalidArgument as 400 in vLLM chat processor - #15902

Open
flpanbin wants to merge 2 commits into
ai-dynamo:mainfrom
flpanbin:fix/vllm-processor-invalid-argument-500
Open

flpanbin wants to merge 2 commits into
ai-dynamo:mainfrom
flpanbin:fix/vllm-processor-invalid-argument-500

Conversation

@flpanbin

@flpanbin flpanbin commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Follow-up to #15305. With --dyn-chat-processor vllm, a worker that rejects a request with a typed InvalidArgument reaches the Python chat processor as a generic error, and the client gets an opaque 500 Internal server error instead of a 400 carrying the worker's message. Two layers drop the error's identity:

  • Bindings: RoutedEngine.generate() flattened failures via to_pyerr, keeping only the Display text and discarding the source chain — so Python never receives a typed exception.
  • Frontend: VllmProcessor's generic except Exception relies on a BackendInvalidArgument: prefix check that matches nothing since refactor(errors): unify semantic error classification #14395 changed the error Display to the normalized class name.

Changes

  • lib/bindings/python/rust/errors.rs — new semantic_to_pyerr: walk the source chain; if a classified DynamoError is InvalidRequest with an explicitly public message, raise typed Python InvalidArgument with that message.
  • lib/bindings/python/rust/llm/routed_engine.rs — use it at the RoutedEngine.generate() dispatch boundary.
  • components/src/dynamo/frontend/vllm_processor.py — re-raise typed InvalidArgument ahead of the generic handler.
  • Tests: add 2 new Rust cases in errors.rs and test_vllm_processor_unit.py‎
  • update comments: lib/llm/src/kv_router/prefill_router/mod.rs, components/src/dynamo/vllm/handlers.py.

Validation

Unit tests:

cargo test --manifest-path lib/bindings/python/Cargo.toml errors::tests   # 4 passed
python3 -m pytest components/src/dynamo/frontend/tests/test_vllm_processor_unit.py \
  -k "invalid_argument or backend_rejection" -x -q                        # 2 passed

E2E:
using a multimodal request against multimodal-disabled workers as the rejection scenario (frontend with--dyn-chat-processor vllm):

curl -s -X POST http://<frontend>:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"<model>","messages":[{"role":"user","content":[{"type":"text","text":"What is this?"},{"type":"image_url","image_url":{"url":"data:image/png;base64,..."}}]}],"max_tokens":20}'
  • Before: 500 {"message":"Internal server error",...}
  • After: 400 {"message":"Received multimodal data but multimodal processing is not enabled. Use --enable-multimodal flag to enable multimodal processing.","type":"Bad Request","code":400}

Related Issues

⚠️ This section is required. Every pull request references the issue it implements.

Summary by CodeRabbit

  • Bug Fixes
    • Routed generation now preserves actionable invalid-request messages instead of returning a generic internal error.
    • Errors without a public message continue to receive a generic response, helping prevent internal details from being exposed.
    • Improved handling of errors passed through the frontend so relevant messages can reach clients.

…rocessor

Signed-off-by: bin <bin.pan@daocloud.io>
@copy-pr-bot

copy-pr-bot Bot commented Oct 8, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@flpanbin
flpanbin deployed to external_collaborator October 8, 2026 15:24 — with GitHub Actions Active
@flpanbin
flpanbin deployed to external_collaborator October 8, 2026 15:24 — with GitHub Actions Active
@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

👋 Hi flpanbin! Thank you for contributing to ai-dynamo/dynamo.

Just a reminder: The NVIDIA Test Github Validation CI runs an essential subset of the testing framework to quickly catch errors.Your PR reviewers may elect to test the changes comprehensively before approving your changes.

🚀

@github-actions github-actions Bot added fix external-contribution Pull request is from an external contributor backend::vllm Relates to the vllm backend frontend `python -m dynamo.frontend` and `dynamo-run in=http|text|grpc` labels Oct 8, 2026
@pull-request-size pull-request-size Bot removed the size/S label Oct 9, 2026
@flpanbin
flpanbin deployed to external_collaborator October 9, 2026 08:08 — with GitHub Actions Active
@github-actions github-actions Bot added the router Relates to routing, KV-aware routing, etc. label Oct 9, 2026
@flpanbin
flpanbin force-pushed the fix/vllm-processor-invalid-argument-500 branch from b4b2b3a to 99ab0d0 Compare October 9, 2026 13:58
@flpanbin
flpanbin deployed to external_collaborator October 9, 2026 13:58 — with GitHub Actions Active
@flpanbin
flpanbin force-pushed the fix/vllm-processor-invalid-argument-500 branch from 99ab0d0 to ea14084 Compare October 9, 2026 14:49
@pull-request-size pull-request-size Bot removed the size/L label Oct 9, 2026
@flpanbin
flpanbin deployed to external_collaborator October 9, 2026 14:49 — with GitHub Actions Active
@flpanbin
flpanbin force-pushed the fix/vllm-processor-invalid-argument-500 branch from ea14084 to 5f6de53 Compare October 9, 2026 15:11
@flpanbin
flpanbin deployed to external_collaborator October 9, 2026 15:11 — with GitHub Actions Active
@flpanbin
flpanbin marked this pull request as ready for review October 9, 2026 15:23
@flpanbin
flpanbin requested review from a team as code owners October 9, 2026 15:23
@coderabbitai

coderabbitai Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: ai-dynamo/dynamo/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: da8120ef-30a8-4419-9612-a3ff294b55a3

📥 Commits

Reviewing files that changed from the base of the PR and between 8120f69 and 5f6de53.


📒 Files selected for processing (6)
  • components/src/dynamo/frontend/tests/test_vllm_processor_unit.py
  • components/src/dynamo/frontend/vllm_processor.py
  • components/src/dynamo/vllm/handlers.py
  • lib/bindings/python/rust/errors.rs
  • lib/bindings/python/rust/llm/routed_engine.rs
  • lib/llm/src/kv_router/prefill_router/mod.rs

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.



Walkthrough

The routed engine now maps classified invalid-request errors to Python InvalidArgument exceptions when a public message is present. The vLLM chat processor preserves these exceptions. A regression test checks propagation of the exception and its message.

Changes

Routed error propagation

Layer / File(s) Summary
Map routed errors to Python exceptions
lib/bindings/python/rust/errors.rs, lib/bindings/python/rust/llm/routed_engine.rs, lib/llm/src/kv_router/prefill_router/mod.rs
The routed engine converts a classified InvalidRequest with a public message to Python InvalidArgument. Other errors become a generic Python Exception. A test checks message lookup through an error source chain.
Preserve invalid arguments in the chat processor
components/src/dynamo/frontend/vllm_processor.py, components/src/dynamo/frontend/tests/test_vllm_processor_unit.py, components/src/dynamo/vllm/handlers.py
The processor re-raises InvalidArgument unchanged. A regression test checks that the exception and message propagate. The handler comment describes which exception messages reach the client.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~20 minutes

Severity of issue fixed: Medium

Merge Risk: ⚪ Minimal · up to 5f6de

Worker-side invalid-argument rejections now reach clients as HTTP 400 with the worker's message. No merge-blocking risk was identified.

Pre-merge checks | Passed 4 | Failed 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage Warning Docstring coverage is 46.67% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 15 functions across 6 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check Passed Issue #15990 requires worker-side typed InvalidArgument rejections to reach the client as HTTP 400 responses with the rejection reason when --dyn-chat-processor vllm is active. semantic_to_pyerr…
Out of Scope Changes check Passed The changes stay within issue #15990. The Rust conversion, routed-engine integration, frontend exception handling, and regression tests directly implement the required error path. The updates in `hand…
Description check Passed The description clearly explains the problem, implementation, validation, and related issue. It includes the required Related Issues section with a concrete issue reference. The required "Where should…
Title check Passed The title clearly and concisely describes the main change: propagating worker-side InvalidArgument errors as HTTP 400 responses in the vLLM chat processor.


  • Fix all pre-merge checks with AI
  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

Comment thread lib/bindings/python/rust/errors.rs Outdated
}
}

/// Walks the error chain for an `InvalidRequest` and returns

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This private helper's doc comment mostly restates the loop and return value, and the only lasting security point about not promoting diagnostic text is already documented immediately above for the same conversion path. Remove the redundant narration rather than adding another copy of the rule.

🤖 AI Fix

Delete this three-line doc comment.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done.

… dispatch

Signed-off-by: bin <bin.pan@daocloud.io>
@flpanbin
flpanbin force-pushed the fix/vllm-processor-invalid-argument-500 branch from 5f6de53 to d353d86 Compare October 10, 2026 01:36
@flpanbin
flpanbin deployed to external_collaborator October 10, 2026 01:36 — with GitHub Actions Active

This branch was successfully deployed

1 active deployment
external_collaborator — d353d86e Deployed Oct 10, 2026 by flpanbin via ok-to-test #29773
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backend::vllm Relates to the vllm backend external-contribution Pull request is from an external contributor fix frontend `python -m dynamo.frontend` and `dynamo-run in=http|text|grpc` router Relates to routing, KV-aware routing, etc. size/M

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG]: Worker-side rejections surface as 500 instead of 400 with --dyn-chat-processor vllm

1 participant