fix(agent): prevent text-only main model from leaking into vision provider client (#57948) - #60991
Closed
webtecnica wants to merge 1 commit into
Closed
webtecnica wants to merge 1 commit into
webtecnica wants to merge 1 commit into
Conversation
Collaborator
|
Thanks for the focused investigation. Current
Automated hermes-sweeper review. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes a bug where
vision_analyzereturns a 400 "Unexpected item type in content" on its first call when the main model is text-only (e.g.deepseek-chat) and no vision auxiliary model is explicitly configured. The error occurs because the text-only main model leaks into the vision provider's client initialization.Root Cause
In
resolve_provider_client(), whenmodelisNoneandprovideris not"auto", the fallback logic calls:This unconditionally falls back to
_read_main_model()— the main runtime model — even when the client is being created for a vision task. If the main model is text-only, it gets injected as the model name for the vision provider client, causing the upstream API to reject the request with a 400 error on the very first message (since a text-only model receives image content it cannot process).Change
In
agent/auxiliary_client.py, the vision-aware fallback now:_get_aux_model_for_provider(provider)if available (same as before)._read_main_model()only whenis_visionisFalse— vision tasks skip this fallback entirely.Verification
resolve_provider_client→ fallback model assignment for vision tasksis_visionguard is specific to this single entry point