Skip to content

fix(bedrock): stop sending output_config.format on AnthropicBedrock aux routes; probe with maxTokens=16 - #124931

Open
dskwe wants to merge 2 commits into
NousResearch:mainfrom
dskwe:fix/bedrock-titlegen-probe-124923
Open

dskwe wants to merge 2 commits into
NousResearch:mainfrom
dskwe:fix/bedrock-titlegen-probe-124923

Conversation

@dskwe

@dskwe dskwe commented Sep 27, 2026 •

Copy link
Copy Markdown

Summary

Two Bedrock fixes from #124923:

  1. output_config.format sent to AnthropicBedrock (400 on every structured-output aux call). _AnthropicCompletionsAdapter.create translated response_format (top-level and extra_body) into Anthropic output_config.format even when the wrapped client is AnthropicBedrock. Bedrock's InvokeModel endpoint rejects that field — 400 {'message': 'output_config.format: Extra inputs are not permitted'} (ValidationException on InvokeModelWithResponseStream for global.anthropic.claude-opus-5-5, claude-opus-4-8, claude-sonnet-5 — endpoint, not model). The _REJECTED_ROUTES memo is per-process, so every new gateway/CLI process re-learns it by eating a 400 on each new session's title call. The adapter now detects AnthropicBedrock clients (class-name check — the anthropic SDK is an optional extra) and skips the translation up front; schema enforcement degrades to prompt compliance, the same fallback used for providers that don't support a format at all.

    I did not take the issue's alternative of unsupported_response_formats on the bedrock provider profile: that filter runs in _build_call_kwargs before the wire is known, and it would also strip response_format from the Mantle route (Bedrock-hosted OpenAI models, a plain OpenAI client with no evidence of rejection), losing structured output there for no reason. The per-client gate scopes the skip to the endpoint that actually rejects.

  2. Context probe maxTokens=8 is below the OpenAI-on-Bedrock minimum. probe_bedrock_context_length sent inferenceConfig={"maxTokens": 8}; OpenAI-on-Bedrock models (openai.gpt-6-sol, openai.gpt-6-luna) reject max output < 16 with a ValidationException before the prompt-length check, so the probe never got the "prompt is too long" error it parses — nothing cached, the two probe tiers (~1.3M/2.2M tokens) re-ran in every new process, and the model resolved to BEDROCK_DEFAULT_CONTEXT_LENGTH (128K), triggering early compression. The probe now sends maxTokens=16.

Root cause

  • agent/auxiliary_client.py: _AnthropicCompletionsAdapter.__init__ now records _is_bedrock_client = type(real_client).__name__ == "AnthropicBedrock"; create() routes response_format translation through a gate that logs-and-skips for Bedrock clients (both the top-level and extra_body forms).
  • agent/bedrock_adapter.py: probe_bedrock_context_length sends maxTokens=16 with a comment explaining the OpenAI-on-Bedrock minimum.

How I tested it

  • New regression tests:
    • tests/agent/test_auxiliary_client.py::TestAnthropicAuxiliaryReasoningTranslation::test_bedrock_client_never_receives_output_config_format — an AnthropicBedrock-named client never gets output_config.format on the wire.
    • ...::test_non_bedrock_client_still_gets_output_config_format — the skip is scoped to Bedrock; plain Messages clients keep the translation.
    • tests/agent/test_bedrock_adapter.py::TestBedrockContextProbe::test_probe_max_tokens_meets_openai_bedrock_minimum — the probe sends maxTokens >= 16 and parses the limit.
  • Full local runs (venv pytest): tests/agent/test_bedrock_adapter.py + tests/agent/test_bedrock_context_cache.py → 135 passed; tests/agent/test_auxiliary_client.py + tests/agent/test_model_metadata.py → 331 passed.

Fixes #124923

…ux routes; probe with maxTokens=16

title_generation and other structured-output aux calls translate
response_format to output_config.format even for AnthropicBedrock clients,
but Bedrock's InvokeModel endpoint rejects that field with 400
'output_config.format: Extra inputs are not permitted' (every Claude
model). The per-process rejection memo forgets this on every restart, so
each new session's title call re-hits the 400 and falls through the retry
ladder. Skip the translation up front for AnthropicBedrock clients; schema
enforcement degrades to prompt compliance.

The Bedrock context-length probe sent maxTokens=8, below the
OpenAI-on-Bedrock minimum of 16: gpt-6-sol/luna raise a ValidationException
(max_output_tokens below minimum) before the prompt-length check, so the
probe never parsed a limit — it burned ~3.5M tokens of probe requests per
process and resolved gpt-6 models to the 128K default, triggering early
compression. Probe with 16.

Fixes NousResearch#124923
…and the 16-token probe floor

- adapter: AnthropicBedrock clients must never receive output_config.format
  (Bedrock 400 'Extra inputs are not permitted', re-learned every process);
  non-Bedrock Messages clients must still get the translation.
- probe: inferenceConfig.maxTokens >= 16 so OpenAI-on-Bedrock models reach
  the prompt-length check instead of failing validation first.
@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint provider/bedrock AWS Bedrock (boto3, IAM) provider/openai OpenAI / Codex Responses API labels Sep 27, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P2 Medium — degraded but workaround exists provider/bedrock AWS Bedrock (boto3, IAM) provider/openai OpenAI / Codex Responses API type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Bedrock: title_generation sends output_config.format (rejected); context probe maxTokens=8 fails on OpenAI-on-Bedrock models

2 participants