Repository navigation
Conversation
…ux routes; probe with maxTokens=16 title_generation and other structured-output aux calls translate response_format to output_config.format even for AnthropicBedrock clients, but Bedrock's InvokeModel endpoint rejects that field with 400 'output_config.format: Extra inputs are not permitted' (every Claude model). The per-process rejection memo forgets this on every restart, so each new session's title call re-hits the 400 and falls through the retry ladder. Skip the translation up front for AnthropicBedrock clients; schema enforcement degrades to prompt compliance. The Bedrock context-length probe sent maxTokens=8, below the OpenAI-on-Bedrock minimum of 16: gpt-6-sol/luna raise a ValidationException (max_output_tokens below minimum) before the prompt-length check, so the probe never parsed a limit — it burned ~3.5M tokens of probe requests per process and resolved gpt-6 models to the 128K default, triggering early compression. Probe with 16. Fixes NousResearch#124923
…and the 16-token probe floor - adapter: AnthropicBedrock clients must never receive output_config.format (Bedrock 400 'Extra inputs are not permitted', re-learned every process); non-Bedrock Messages clients must still get the translation. - probe: inferenceConfig.maxTokens >= 16 so OpenAI-on-Bedrock models reach the prompt-length check instead of failing validation first.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Two Bedrock fixes from #124923:
output_config.formatsent to AnthropicBedrock (400 on every structured-output aux call)._AnthropicCompletionsAdapter.createtranslatedresponse_format(top-level andextra_body) into Anthropicoutput_config.formateven when the wrapped client isAnthropicBedrock. Bedrock's InvokeModel endpoint rejects that field —400 {'message': 'output_config.format: Extra inputs are not permitted'}(ValidationException on InvokeModelWithResponseStream forglobal.anthropic.claude-opus-5-5,claude-opus-4-8,claude-sonnet-5— endpoint, not model). The_REJECTED_ROUTESmemo is per-process, so every new gateway/CLI process re-learns it by eating a 400 on each new session's title call. The adapter now detectsAnthropicBedrockclients (class-name check — the anthropic SDK is an optional extra) and skips the translation up front; schema enforcement degrades to prompt compliance, the same fallback used for providers that don't support a format at all.I did not take the issue's alternative of
unsupported_response_formatson the bedrock provider profile: that filter runs in_build_call_kwargsbefore the wire is known, and it would also stripresponse_formatfrom the Mantle route (Bedrock-hosted OpenAI models, a plain OpenAI client with no evidence of rejection), losing structured output there for no reason. The per-client gate scopes the skip to the endpoint that actually rejects.Context probe
maxTokens=8is below the OpenAI-on-Bedrock minimum.probe_bedrock_context_lengthsentinferenceConfig={"maxTokens": 8}; OpenAI-on-Bedrock models (openai.gpt-6-sol,openai.gpt-6-luna) reject max output < 16 with a ValidationException before the prompt-length check, so the probe never got the "prompt is too long" error it parses — nothing cached, the two probe tiers (~1.3M/2.2M tokens) re-ran in every new process, and the model resolved toBEDROCK_DEFAULT_CONTEXT_LENGTH(128K), triggering early compression. The probe now sendsmaxTokens=16.Root cause
agent/auxiliary_client.py:_AnthropicCompletionsAdapter.__init__now records_is_bedrock_client = type(real_client).__name__ == "AnthropicBedrock";create()routesresponse_formattranslation through a gate that logs-and-skips for Bedrock clients (both the top-level andextra_bodyforms).agent/bedrock_adapter.py:probe_bedrock_context_lengthsendsmaxTokens=16with a comment explaining the OpenAI-on-Bedrock minimum.How I tested it
tests/agent/test_auxiliary_client.py::TestAnthropicAuxiliaryReasoningTranslation::test_bedrock_client_never_receives_output_config_format— an AnthropicBedrock-named client never getsoutput_config.formaton the wire....::test_non_bedrock_client_still_gets_output_config_format— the skip is scoped to Bedrock; plain Messages clients keep the translation.tests/agent/test_bedrock_adapter.py::TestBedrockContextProbe::test_probe_max_tokens_meets_openai_bedrock_minimum— the probe sendsmaxTokens >= 16and parses the limit.tests/agent/test_bedrock_adapter.py+tests/agent/test_bedrock_context_cache.py→ 135 passed;tests/agent/test_auxiliary_client.py+tests/agent/test_model_metadata.py→ 331 passed.Fixes #124923